HN user

Panos

2,703 karma
Posts38
Comments76
View on HN
www.behind-the-enemy-lines.com 15h ago

I used ChatGPT to sue a Norwegian airline from New York and get $4760

Panos
30pts11
www.behind-the-enemy-lines.com 6mo ago

Fighting Fire with Fire: Scalable Oral Exams with an ElevenLabs Voice AI Agent

Panos
5pts0
medium.com 4y ago

Pricing Homes Like Agents Do: “Human-in-Charge AI” for Pricing Using Comparables

Panos
8pts1
stackoverflow.com 10y ago

Why is char[] preferred over String for passwords in Java?

Panos
2pts0
www.behind-the-enemy-lines.com 10y ago

The Decline of Mechanical Turk: A Cohort Analysis of MTurk Requesters

Panos
4pts0
www.behind-the-enemy-lines.com 12y ago

My Peer Grading Scheme

Panos
2pts0
john-joseph-horton.com 12y ago

Tools that developers are using: Data from the monitoring tool of oDesk

Panos
1pts0
www.behind-the-enemy-lines.com 12y ago

Online labor markets: Why they can't scale and the crowdsourcing solution.

Panos
2pts0
news.ycombinator.com 13y ago

I'd post an NP-complete joke but if you've heard one, you've heard them all.

Panos
1pts0
www.behind-the-enemy-lines.com 13y ago

Identity verification and marketplaces: Why true anonymity is not an option

Panos
2pts0
www.behind-the-enemy-lines.com 13y ago

Intrade Archive: Data for Posterity

Panos
1pts0
www.behind-the-enemy-lines.com 13y ago

WikiSynonyms: Find synonyms using Wikipedia redirects

Panos
1pts0
www.behind-the-enemy-lines.com 13y ago

Is Mechanical Turk a $10 billion dollar company?

Panos
1pts0
realgl.blogspot.com 13y ago

The "After" Image: Crowdsourcing Cartoon Generation

Panos
2pts0
www.behind-the-enemy-lines.com 13y ago

The disintermediation of the firm: The feature belongs to individuals

Panos
1pts0
www.behind-the-enemy-lines.com 14y ago

The oDesk Flower: Playing with Visualizations

Panos
4pts0
www.behind-the-enemy-lines.com 14y ago

How I attacked myself using Google and I ramped up a $1000 bandwidth bill

Panos
775pts142
www.behind-the-enemy-lines.com 14y ago

Philippines: The country that never sleeps

Panos
53pts20
www.behind-the-enemy-lines.com 14y ago

50% of the online ads are never seen

Panos
9pts1
www.behind-the-enemy-lines.com 14y ago

Crowdsourcing and the end of job interviews

Panos
7pts2
www.behind-the-enemy-lines.com 14y ago

The Need for Standardization in Crowdsourcing

Panos
12pts1
onlinelabor.blogspot.com 14y ago

High-wage skills, or why you might want to learn Clojure if you're not a lawyer

Panos
82pts62
www.behind-the-enemy-lines.com 14y ago

Mechanical Turk vs oDesk

Panos
51pts16
behind-the-enemy-lines.blogspot.com 15y ago

Why I will never pursue cheating again

Panos
691pts482
behind-the-enemy-lines.blogspot.com 15y ago

Crowdsourcing, Linked Data, Humans, and Machines: Rediscovering a Lost Treasure

Panos
44pts4
behind-the-enemy-lines.blogspot.com 15y ago

Crowdsourcing Education

Panos
1pts0
www.economist.com 15y ago

How spammers and robots exploit human workers on Mechanical Turk

Panos
6pts0
www.slate.com 15y ago

Awsum Shoes! Is it ethical to fix grammatical errors in Internet reviews?

Panos
6pts1
behind-the-enemy-lines.blogspot.com 15y ago

Pay Enough or Don't Pay at All

Panos
129pts20
www.slideshare.net 15y ago

Crowdsourcing: Lessons from Henry Ford

Panos
2pts0

OP here. I wrote this up because the discourse around AI in legal contexts usually swings between "it replaces everyone" and "it hallucinates case law and gets you sanctioned." This was the case in the middle, where it actually shone: navigating an international dispute in a foreign conciliation court (the Norwegian Forliksråd), where even finding a human lawyer with all the necessary knowledge would be challenging, and paying one would be completely inefficient.

As I note at the end, the AI eliminated the knowledge bottleneck, but it still took 11 months of waiting for companies, agencies, and courts to reply.

Happy to answer questions about the routing or the EU261 technicalities. And to preempt the obvious one: a credit card chargeback only covers the original ticket cost. This was about getting the statutory €600/passenger penalty plus our overnight care expenses (and it was mainly the latter that prompted the whole story; Norse would probably have saved money if it had just arranged hotel rooms instead of handing us a $25 voucher).

Point well taken. :-)

I actually touch on this exact risk at the very end of the post. AI can automate drafting and knowledge retrieval, but that is only a fraction of the overall process.

The legal system is fundamentally slow, by design. It still took 11 months of waiting for agencies, airlines, and courts to move. That inherent slowness is creating friction to filter out low-effort "slop"; that latency partially accomplishes that.

Not an issue of cost, at all.

Absolutely the easiest solution would have been to have a written exam on the cases and concepts that we discussed in class. It would take a few hours to create and grade the exam.

But at a university you should experiment and learn. What better class to experiment and learn than the “AI Product Management”. Students were actually intrigued by the idea themselves.

The key goal: we wanted to ensure that the projects that students submitted was actually their own work, not “outsourced” (in a general sense) to teammates or to an LLM.

Gemini 3 and NotebookLM with slide generation were released in the middle of the class, and we realized that it is feasible for a student to have a flaweless presentation in front of the class, without understanding deeply what they are presenting.

We could schedule oral exams during the finals week, which would be a major disruption for the students, or schedule exams during the break, violating university rules and ruining students vacation.

But as I said, we learned that AI-driven interviews are more structured and better than human-driven ones, because humans do get tired, and they do have biases based on who is the person they are interviewing. That’s why we decided to experiment with voice AI for running the oral exam.

I agree that I am not yet confident to use this approach for my technical classes. I am still very unhappy with any option for assessment for technical classes, but I would not trust an LLM to come up with good questions. NotebooksLM does come up with decent quizzes, but nothing super hard.

For the use of LLM in classes: I understand the reasoning, but I found LLMs to be extremely educational for parsing through dense material (eg parsing an NTSB report for an Uber self-driving crash). Prohibiting students from using LLMs would be counterproductive.

But I still want students to use LLMs responsibly, hence the oral exam.

By the way the voice agent flagged the system as “the student is obviously fooling around”. I was expecting this to be caught during the grading phase but ElevenLabs has done such a good work with their product.

Guys, thank you for such fooling around. All these adversarial discussions will be great for stress testing the system. Very likely we will use these conversations as part of the course in the Spring to get students to see what it means to let AI systems “in the wild”.

Not the case for the class in the blog post, but we also have many online classes. Many professionals prefer these online classes because they can attend without having to commute, and can do it from a place of their own convenience.

Such classes do not have the luxury of pen-and-paper exams, and asking people to go to testing centers is a huge overkill.

Take home exams for such settings (or any other form of written exam) are becoming very prone to cheating, just because the bar to cheating is very low. Oral exams like that make it a bit harder to cheat. Not impossible, but harder.

Just in case, I am the author of the blog post. For our "AI" class, it felt like a good class to experiment with something novel.

No, we do not want to eliminate the pen and paper exam. It works well. We use it.

The oral exam is yet another tool. Not a solution for everything.

In our case, we wanted to ensure that the students who worked on the team project: (a) contributed enough to understand the project, (b) actually understood their own project and did not rely solely on an LLM. (We do allow them to use LLMs, it would be stupid not to.)

The students who did badly in the oral exam were exactly the students who we expected to do badly in the exam, even though they aced their (team) project presentations.

Could we do it in person? Sure, we could schedule personalized interviews for all the 36 students. With two instructors, it would have taken us a couple of days to go through. Not a huge deal. At 100 students and one instructor, we would have a problem doing that.

But the key reason was the following: research has shown that human interviewers are actually worse when they get tired, and that AI is actually better for conducting more standardized and more fair interviews. That result was a major reason for us to trust a final exam on a voice agent.

Stripe Atlas 10 years ago

This is a regulatory requirement, part of the "Know Your Customer" doctrine. In plain words, banks are required to know who is the client who has opened an account.

Most banks will be risk averse and will not open an account to anyone applying online from abroad. Even for US persons applying online, they will ask quite a bit of documentation.

Some banks (but not all) will open an account for a non-US person, when the non-US person physically visits a US branch, with proper identification (typically a passport) and documentation on why they want the account. But even in such cases, it is up to the discretion of the bank employee to decide whether the risk of opening an account for a non-US person is worth the benefit. So, the same bank may give different replies to the same inquiry, depending on the branch asked.

As a concrete example, TD Ameritrade will open easily an account for a foreigner in the Chinatown branch in NYC, but will not open an account when the same customer visits a branch in midtown in NYC.

The comparison with Uber and Whatsapp is not the proper one. These are private companies that were funded and acquired, respectively, purely on growth potential.

OpenTable has been a public company for almost 5 years now (see http://finance.yahoo.com/echarts?s=OPEN). Revenues, cost, growth, and all other metrics have been publicly examined and scrutinized for long time. The 46% premium paid by Priceline is based on how the new management estimates that they can leverage the assets of Opentable and hardly a "bubble-ish" premium.

If you believe that OpenTable is part of a bubble, then the whole US stock market is in a bubble, which may be true but again not directly connected to Uber and Whatsapp valuations.

Thanks for the offer. I have a decently long experience with MTurk to know how to get things done.

My point is that microtask work is not necessarily the optimal setting for tasks that are expected to last for longer periods of time. It is often beneficial to train and give people meaningful pieces of work instead of converting real work into micro-work and assume that workers are not intelligent enough to get things done properly.

All my work on Mechanical Turk is to feed the data into a machine learning process that automates the task. Some tasks are easy to automate (any classification task), some others are harder (anything related to content generation, or vision).

I have to say though that there is a pattern: Once you solve and automate a process, people want more, and push you into doing a task that cannot be easily automated. Then you try to automate it, and the cycle continues...

Let me suggest one solution for such cases: Use iterative tasks, in the spirit introduced by TurkIt.

For example, you want to create a caption for an image. You let a user create a caption. Then you take this caption and give it to another user, asking the user to improve it. Take the two versions and ask other workers, "which of the two versions is better?". Iterate until no improvement is possible.

Not a trivial setup, but gets around the binary accept/reject decisions problem and generates results of significantly superior quality.

Yes, in retrospect, we can easily agree that you are correct.

I will just give here the answer that I posted in another thread http://hackerne.ws/item?id=2797371

the post was supposed to be "a story with the twist." Had I known that I was going to have hundreds of thousands of people reading the post, I would have followed the standard journalistic practice of writing a summary at the very first section. (See http://t.co/2kJEkJW for a copy of the post. Note: I asked the post to be taken down until I repost the original article but the journalist is really playing childish games.)

I kind of felt this a few hours after the post went out, so I added two clarification points early on in the article:

1. I am not giving up the fight, I will just fight differently, please see the conclusions (link),

2. This was not about NYU and people cheating in business schools; people cheat everywhere: the story gives an explanation why they remain undetected.

Oh well, people could not even read these two points.

But I want to write stories in my blog, not papers with an abstract, executive summary and table of contents.

Yes, I believe that there is a 1%-2% of blatant cheater everywhere, and a significantly bigger percentage of students that take liberties with their homework assignments. (Asking people for help, looking at solutions, collaborating in individual assignments, etc). There is no clear black and white separation of "cheaters" vs "honest" but a full spectrum of gray.

Regarding the parable, I see the point and how what I wrote could be interpreted in the way that you describe. My intention was to refer to the city as the whole academic system. The "Redwich Village" is a reference to Greenwich Village, the neighborhood in NYC where NYU is located.

So, Redwich Village stands for my institution, which was attacked ferociously as being the harbor of cheaters over the last few days, while people fail to recognize that this is most probably a more generic problem. You may claim that I overgeneralize without proof, but I have sufficient evidence from other people that this is the case. Especially after all the emails that I received within the last few days where people described their own experiences with cheating that were very similar (most did so anonymously, but with specific names of top universities around the world).