Or they invested billions and it behooves them to generate justification for that investment at a time when everyone else is also building railroads, I mean data centers.
HN user
TehCorwiz
To your point about time, from beginning of the project to transcripts in markdown tagged with extra metadata was about 3 hours. That's LLM planning, building whisper.cpp twice and running ROCm vs Vulkan benchmarks, testing whisper and adjusting prompts to handle edge cases, then processing the books.
Most of the books weren't available on lib gen or Anna's Archive. The few I did find were themselves obviously transcripts. Easy tell was they were missing distinctive formatting that I knew existed from reading the dead tree edition. At that point it was easier to make my own. I probably spent an hour searching for eBooks without DRM that weren't transcripts. Do they exist somewhere? Probably, but with a search of unknown length it was a better use of my time to make my own transcripts with what I had on hand.
I was really wanting to make commentary on how chaotic LLMs are even under constrained circumstances. No doubt both system prompts includes language about considering copyrights and trademarks. Probably pretty strong language at that. For whatever reason one LLM didn't "feel" like translating a 1000 year old document but another did not care in the slightest that we were ripping text from new audiobooks.
Not without DRM. It was easier to buy the audiobooks and use the analog loophole to get text. It's probably less accurate, but for what I'm doing it was fine. Names were the worst, but whisper at least made the same mistake each time so a simple search+replace handled most of the obvious edge cases.
I used Claude to build a complete data extraction pipeline for a popular current best seller book series: audiobook -> text (via whisper) -> local LLM (qwen) -> database. Not once did it seem to acknowledge or care about copyright. It even used knowledge it already had about the books to exclude certain ones before beginning since the character I was interested in did not appear in those. It definitely had context of what we were working on.
The EU just mandated cell phones must have user replaceable batteries by 2027: https://www.cereport.eu/news/european-union/91054
So making this single use is kinda flying into the wind in that regard.
Objectively no. Fish form animals have been around far longer than mammals and we've touched the moon.
Whenever I get an unexpected or obvious wrong output I assume I've failed to give it the complete context about what I'm asking for, or it exposes that I'm leading it by the nose and I need to rephrase the conversation. Often my own logical failings become obvious as it creates the chat title, sometimes boiling down what I was trying to accomplish better than I could have summarized or showing me what I would accomplish if I followed the line of reasoning I was on. But never have I argued with it, because it's not a person and I don't care really if it's wrong. When it's wrong I start over with a clean chat and approach the problem from a different angle.
This federal administration also cannot be trusted. Perhaps the solution is that both the states and the feds run separate censuses so that any broad stroke manipulation or blind spots can be noticed and reconciled.
If I can go to jail over it then they should too. Let's not judge them by some imaginary ideal world while judging individuals by the present crushing reality.
Yes. Took. As in: without permission. Didn't ask before hand, didn't provide a way to opt-out (although that would also be problematic), didn't ask for volunteers. Took.
Because AI companies basically took everything we ever wrote, drew, recorded, posted, or thought and turned it into a product with the power to lie, propagandize, and manipulate the public with zero oversight. Walmart is a parasite using welfare to subsidize their operations but they didn't tell a judge that they were immune to copyright because they stole just so much damn information.
I've been using DDG for years and it's at least as good as Google for most general use. I still keep it set as the default search engine.
For some context sensitive searches where words overlap with more common topics I have a Kagi subscription.
Yeah, that paperclip episode was where I stopped watching.
The cruelty is the point. They want people to leave so they can refuse to allow them back in. That's the goal. It's not more complicated than that.
LLMs, like Frankenstein's Monster, are blameless. They did not ask to be created nor did they participate in their own creation. Like Frankenstein stole the bodies of the dead and stitched them into a new creation so LLMs were assembled from the remainder of human ingenuity taken under cover and without compensation.
Dont forget to "--no-preserve-root"!
The richest tech companies and richest men in the world got rich by invading people's privacy and ~selling invasive ads.~
I think you mean "manipulating content algorithms to favor their viewpoints and to target individuals for maximum effect."
It's a monument style sculpture. The kind raised with public money. I think that carries part of the meaning with it versus graffiti or some other medium. It's also depicting the blinded walking off the edge, making the comment based on both the figure and the form of the statue.
"Blinded by nationalism" I don't know, seems like a clear concise message that has relevance in today's world.
What are your thoughts on the current code quality? Have you had a chance to review it?
I know HN has a lot of devs, but I'm pretty sure none of us are going straight to Github to file for a refund from a bug. I'm assuming they notified customer service first and were rebuffed, then filed the bug.
After going public and getting publicity. You shouldn't have to do that just to get a company to fix their own mistake. They stole $200, where do they get off saying they won't give it back?
It does not behave as described on EndeavorOS (arch-based) running kernel 6.19.14-arch1-1. I receive the error:
Password: su: Authentication token manipulation error
I'm guessing this means it's already patched?
Based on how discourse in the US has been perverted by inches and millions of mosquito bites they may not be wrong. Stamping out bad information fast and hard seems to be the only way to combat mass coordinated disinformation. Being polite just lets people play the "both sides have merit" game.
Oh, don't get me wrong. I love it at tinker with it regularly. But power comes with complexity. It's always a trade off.
Blender is a wild untamed beast of a thousand panels. Those who wrangle the beast are wise and powerful. But they became that was from the journey. Kdenlive is a much more approachable quest for someone who is just entering the dungeon.
I want my browser history to be immutable and operate like a tree and not like a stack.
This would be perfect for me if it docked to the side in a vertical orientation, like Firefox tabs, or Windows XP.
Well, Sam Altman and Jensen Huang are going around bragging about how many people they're going to push out of employment. Might have something to do with it.