HN user

calny

255 karma

building AI systems for law, insurance, and creative firms tech & insurance lawyer twitter @joe_wilbert jwilbert dot com

Posts3
Comments59
View on HN
Tidal AI Policy 23 days ago

I’m also curious! Are you keeping a secret because you don’t want people to try to find work arounds, kind of like prompting LLMs not to say “delve” or “tapestry” in order to make your AI writing sound less like AI? Or is it something else?

Tidal AI Policy 23 days ago

I'm curious about they will apply the part saying "AI-generated music will not be monetizable." What does AI-generated music mean, exactly? What if you make an AI generated bassline but produce the rest of a track by hand? How about an AI vocal? Or a mix of AI stems and your own recordings?

Tidal's terms and conditions (https://tidal.com/terms) say that:

“AI-Generated Content” means any audio content, inclusive of musical works and sound recordings, that is wholly or substantially generated by generative artificial intelligence, with limited or no direct human creative input beyond an initial text prompt or similar instruction. ... You acknowledge that AI detection technology may produce false positives or false negatives.

And:

If you use TIDAL Upload, your Tracks may be scanned for the purpose of identifying whether the content is AI-Generated Content, and to label such content accordingly on the Tidal platform. You acknowledge that such scanning and labeling is performed on a best-efforts basis and that Tidal shall not be liable for any inaccuracies in AI detection or labeling. AI-Generated Content uploaded to Tidal is not eligible for monetization. If you believe your Tracks were erroneously tagged as AI-Generated, you can reach out to support@tidal.com.

I was reading up on the author and saw this interesting bit[0]:

An algorave (from an algorithm and rave) is an event where people dance to music generated from algorithms, often using live coding techniques. Alex McLean of Slub and Nick Collins coined the word "algorave" in 2011, and the first event under such a name was organised in London, England. It has since become a movement, with algoraves taking place around the world.

[0] https://en.wikipedia.org/wiki/Algorave

You're right, I've followed the litigation closely. I've advocated for years that "training is fair use" and I'm generally an anti-IP hawk who DEFENDS copyright/trademark cases. Only recently have I started to concede the issue might have more nuance than "all training is fair use, hard stop." And I still think Judge Alsup got it right.

That said, even if model training is fair use, model output can still be infringing. There would be a strong case, for example, if the end user guides the LLM to create works in a way that copies another work or mimics an author or artist's style. This case clearly isn't that. On the similarity at issue here, I haven't personally compared. I hope you're right.

The maintainer's response: https://github.com/chardet/chardet/issues/327#issuecomment-4...

The second part here is problematic, but fascinating: "I then started in an empty repository with no access to the old source tree, and explicitly instructed Claude not to base anything on LGPL/GPL-licensed code." Problem - Claude almost certainly was trained on the LGPL/GPL original code. It knows that is how to solve the problem. It's dubious whether Claude can ignore whatever imprints that original code made on its weights. If it COULD do that, that would be a pretty cool innovation in explainable AI. But AFAIK LLMs can't even reliably trace what data influenced the output for a query, see https://iftenney.github.io/projects/tda/, or even fully unlearn a piece of training data.

Is anyone working on this? I'd be very interested to discuss.

Some background - I'm a developer & IP lawyer - my undergrad thesis was "Copyright in the Digital Age" and discussed copyleft & FOSS. Been litigating in federal court since 2010 and training AI models since 2019, and am working on an AI for litigation platform. These are evolving issues in US courts.

BTW if you're on enterprise or a paid API plan, Anthropic indemnifies you if its outputs violate copyright. But if you're on free/pro/max, the terms state that YOU agree to indemnify THEM for copyright violation claims.[0]

[0] https://www.anthropic.com/legal/consumer-terms - see para. 11 ("YOU AGREE TO INDEMNIFY AND HOLD HARMLESS THE ANTHROPIC PARTIES FROM AND AGAINST ANY AND ALL LIABILITIES, CLAIMS, DAMAGES, EXPENSES (INCLUDING REASONABLE ATTORNEYS’ FEES AND COSTS), AND OTHER LOSSES ARISING OUT OF … YOUR ACCESS TO, USE OF, OR ALLEGED USE OF THE SERVICES ….")

Gemini 3.1 Pro 5 months ago

I get it, I just meant the fish is poorly done, when I’d have guessed it would be relatively simple part. Maybe the black dot eye is misplaced idk.

I didn't catch that on first read, but I see why you'd say that. LLMs are ridiculous in the constant usage "it's not X it's Y" -- It's in almost every response from Opus 4.5. "It's not X it's Y" is ruined for regular writing.

I'm also skeptical of anything that claims to reliably detect AI writing. FWIW, I plugged the comment into Pangram Labs, which claims to be the most reliably and seems to have worked well before. It categorized the comment as 100% human written with medium confidence.

Stated more cynically, many platforms have an interest in attention hijacking. Done well, agents' 'laser focused attention' could help users avoid wasting time (wandering attention) and money (impulse buys). This is a good thing, even if it dings revenue of some existing platforms. If a company's business model is impulse buying and ad revenue (this isn't eBay IMO), then good riddance.

This is a really interesting point, and you're right to say it's complicated. I'm sort of an anti-IP hawk (I actually rep defendants in IP cases) and personally agree with it. But the US Copyright Office's position on GenAI supports the opposite view:

Nor do we agree that AI training is inherently transformative because it is like human learning. To begin with, the analogy rests on a faulty premise, as fair use does not excuse all human acts done for the purpose of learning. A student could not rely on fair use to copy all the books at the library to facilitate personal education; rather, they would have to purchase or borrow a copy that was lawfully acquired, typically through a sale or license. Copyright law should not afford greater latitude for copying simply because it is done by a computer. Moreover, AI learning is different from human learning in ways that are material to the copyright analysis. Humans retain only imperfect impressions of the works they have experienced, filtered through their own unique personalities, histories, memories, and worldviews. Generative AI training involves the creation of perfect copies with the ability to analyze works nearly instantaneously. The result is a model that can create at superhuman speed and scale. In the words of Professor Robert Brauneis, “Generative model training transcends the human limitations that underlie the structure of the exclusive rights.”[0]

I disagree with the Copyright Office here, but ofc they're the Copyright Office and I could be wrong. More broadly I'm struggling with how to permit and incentivize creation of powerful generative models while not screwing creators in the process. There are startups and other efforts trying to address this through novel licensing, etc., but AFAIK there's no great solution. I'm also cautiously optimistic that there will be some decentralized and/or federated options. It's complicated indeed.

[0] https://www.copyright.gov/ai/Copyright-and-Artificial-Intell...

PBS Spacetime did an interesting video on DCQE, but it tripped me up trying to fully understand what was happening: https://www.youtube.com/watch?v=8ORLN_KwAgs&t=601s ... Later Sabine Hossenfelder did a video debunking the proposition that DCQE somehow showed that the past was being rewritten. https://www.youtube.com/watch?v=RQv5CVELG3U And Matt from PBS Spacetime acknowledged she was right in this respectful comment:

Sabine, this is amazing. You are, as usual, 100% right. The delayed choice quantum eraser is a prime example of over-mystification of quantum mechanics, even WITHIN the field of quantum mechanics! I (Matt) was guilty of embracing the quantum woo in that episode 5 years ago. Since then I've obsessed over this family of experiments and my thinking shifted quite a bit.

I’m an IP lawyer & AI dev: my first reaction was, “hmm there are trademark issues here.” From a US perspective: “Perplexity” certainly CAN be a trademark, and the company has applied for one—to my knowledge it’s still pending. If the term was merely “descriptive” of the service provided, like “American Airlines”, then the company would need to show that the term has acquired distinctiveness: ie, that purchasers associate the term with that specific company. But perplexity is probably more than merely descriptive here.

Assuming that they have a valid trademark, the issue becomes whether there is a likelihood of confusion between Perplexity and Perplexica. That is a fact-specific, multifactor test, which I’ll spare you. But there could be arguments both ways IMO

EDIT: trademark issues aside, cool project!

Sorry to hear this, but congrats to Bob for a life well lived and building a brand that made quality products. We have their muesli multiple times a week, their farro as well, and this morning our kids loved Valentine's Day pancakes made from their mix. Thanks Bob

I’ll take it. I spend about half my time developing/promptsmithing and the other half lawyering. “Wordsmith” sure beats some of the other lawyer epithets out there

Yep it's Word exported to pdf. Source: Am attorney, do this all the time. You write it up in Word, save as pdf. Then upload it to the court website, which (in federal court, at least) puts the case number in blue text at the top for the officially-filed version.

The 1-28 pleading numbers on the side are annoying. They're specific to courts in California and a few other jurisdictions, and the rules of court require them. But many other courts don't have them, and they only help to cite specific lines within pages; eg "Complaint 5:4-9" means "Complaint at page 5, at lines 4 to 9". It's occasionally useful for court filings like this, but more useful for court/deposition transcripts of testimony to show precisely where a witness said something.

Related: I tried building an RNN to generate legal pleadings back around 2018/19 and gathered a bunch of docs like this from courts across the country as training data. Processing text with those pleading numbers was a pain, so I built a CNN to classify whether a document had pleading numbers or not, which affected downstream processing. Probably the wrong approach in a bunch of ways, but I was just learning.

Very interesting point. It'd be a challenge to execute, but I'd be glad to see a non-profit that lets online communities discuss topics of interest, solve problems, and expand knowledge. Maybe something like that exists and I just don't realize it.

Sorry to make this about AI, but it'd also be interesting whether such a non-profit makes its data fully open--i.e., for AI companies to scoop up--or has more restrictive terms that forbid AI "scooping" without a separate agreement. Lots of tricky issues, trade-offs, and interesting incentives involved. There could be alliances with orgs creating open source models, for example. If anyone is working on a nonprofit like this or just wants to chat about it, please reach out.

Congrats on the launch! There's certainly a "problem" worth solving here. I'm a lawyer, and it's crazy that I occasionally actually have to negotiate non-substantive parts of contracts like severance clauses. There really should be standardized boilerplate for at least some provisions (like you're doing), where companies can quickly say "we are using the common terms", and if another side pushes back it raises questions. That said, it's tricky to deal with edge cases, differences in state law, etc.

Nevertheless, I like your approach. Kudos especially to open sourcing the contracts--in fact, I think that's critical to success in this space. Open sourcing the terms allows companies to understand them and quickly agree to them with confidence. Wishing you luck!