haha brutal. maybe tomorrow
HN user
nonhaver
ill leave it at this: if “zero-hallucination omniscience” is your bar, you’ll stay disappointed - and that’s on your expectations, not the tech. personally i’ve been coding/researching faster and with fewer retries every time a new model drops - so my opinion is based on experience. you’re free to sit out the upgrade cycle
thats hilarious actually. gives credence to the gpt theory haha
not to offend - but it sounds like your response/worries are based more on an emotional reaction. and rightly so, this is by all means a very scary and uncertain time. and undeniably these companies have not taken into account the impact their products will cause and the safety surrounding that.
however, a lot of your claims are false - progress is being made in nearly all the areas you mentioned
hallucinations
are reduced with GPT-5
https://cdn.openai.com/pdf/8124a3ce-ab78-4f06-96eb-49ea29ffb...
"gpt-5-thinking has a hallucination rate 65% smaller than OpenAI o3"
limited context window
same deal. gemini 2.5-pro has a 1 million token context window and GPT-5 is 400k up from 200k with o3
https://blog.google/technology/google-deepmind/gemini-model-...
"native multimodality and a long context window. 2.5 Pro ships today with a 1 million token context window (2 million coming soon)"
expensive to operate and train
we don't know for certain but GPT-5 provides the most intelligence for the cheapest price at $10/1 million output tokens which is unprecedented
https://platform.openai.com/docs/models/gpt-5
guardrails
are very well implemented in certain models like google who provide multiple safety levels
https://ai.google.dev/gemini-api/docs/safety-settings
"You can use these filters to adjust what's appropriate for your use case. For example, if you're building video game dialogue, you may deem it acceptable to allow more content that's rated as Dangerous due to the nature of the game. In addition to the adjustable safety filters, the Gemini API has built-in protections against core harms, such as content that endangers child safety. These types of harm are always blocked and cannot be adjusted."
now id like to ask you for evidence that none of these aspects have been improved - since you claim my examples are vague but make statements like
Inability to recall simple information
inability to stay on task
(doesn't) support its output
(no) long term planning
ive experienced the exact opposite. not 100% of the time but compared to GPT-4 all of these areas have been massively improved. sorry i cant provide every single chat log ive ever had with these models to satisfy your vagueness-o-meter or provide benchmarks which i assume you will brush aside.
as well as the examples ive provided above - you seem to be making claims out of thin air and then claim others are not providing examples up to your standard.
i didnt bring examples because i said personal experience. heres my "evidence" - gpt 4 took multiple shots and iterations and couldnt stay coherent with a prompt longer than 20k tokens (in my experience). then when o4 came out it improved on that (in my experience). o1 took 1-2 shots with less iterations (in my experience). o3 zero shots most of the tasks i throw at it and stays coherent with very long prompts (in my experience).
heres something else to think about. try and tell everybody to go back to using gpt-4. then try and tell people to go back to using o1-full. you likely wont find any takers. its almost like the newer models are improved and generally more useful
you dont remember deepseek introducing reasoning and blowing benchmarks led by private american companies out of the water? with an api that was way cheaper? and then offered the model free in a chat based system online? and you were a big fan?
this is a very odd perspective. as someone who uses LLMs for coding/PRs - every time a new model released my personal experience was that it was a very solid improvement on the previous generation and not just meant to "confuse". the jump from raw GPT-4 2 years ago to o3 full is so unbelievable if you traveled back in time and showed me i wouldn't have thought such technology would exist for 5+ years.
to the point on hallucination - that's just the nature of LLMs (and humans to some extent). without new architectures or fact checking world models in place i don't think that problem will be solved anytime soon. but it seems gpt-5 main selling point is they somehow reduced the hallucination rate by a lot + search helps with grounding.
i think this is more an effect of releasing a model every other month with gradual improvements. if there was no o-series/other thinking models on the market - people would be shocked by this upgrade. the only way to keep up with the market is to release improvements asap
also wondering this. had to pause the livestream to make sure i wasnt crazy. definitely eyebrow raising
there exists tools that can zero shot complex tasks (claude/codex). bar has been raised and jules doesn’t stack up. knowing google it will probably improve in due time
similar experience. i would put codex over claude personally due to the better rate limits (of which i haven’t hit once yet even on extensive days) but jules was not very good - too messy and i prefer alternative outputs to creating a pull request. like in codex you can copy a git patch which is so incredibly useful to add personal tweaks before committing
thats so insane
wild take: critique != censorship. people can consume whatever they want and we can call out when the supply chain is built on unconsented scraping and zero stewardship CamperBob2
Well put and reflects my thoughts exactly. It's borderline concerning there are people who consume this type of media by choice and forethought.
Absolutely horrendous - well and truly. Ignoring the dystopic undertones of a product like this - the audio quality is low and filled with artifacts and destroyed transients. Voices are robotic and generic with basically zero emotion. Lyric generation is frankly embarrassing. Everything about this is horrible. Additionally the UI broke for me multiple times and I couldn't play the track without downloading it.
What is going on at ElevenLabs? Is everything vibe coded now? Is there nobody testing these products before pushing them out? This is the first time I'd characterize a product from a top AI company as complete slop from a conceptual and implementation perspective.
I feel like there's some sort of disconnect between the minds of tech CEOs and the general population's wants/needs especially in regards to creative domains. What kind of human wants this? Is there some grudge of tech people who never learned music/art/etc so their solution is to optimize it and create models so they can feel something?
if im understanding correctly this was a public bucket? aside from the obvious leaking of data couldnt this also be subject to a DoW (denial of wallet) attack where a user could auto download all the images constantly on a VPS and cause a massive bill?
ah man they got us. we thought it was spam when it was a social experiment all along. now look whos got egg on their face.
impressive evals. i wonder how much of that can be attributed to the enhanced context understanding. i feel like that/length are the bottleneck of the majority of commercial models.
not sure why people feel the need to complain in the comments of this anniversary post for a free service. been using the MDN docs for 5+ years and its been an invaluable resource that also promotes exploration - ive stumbled upon so many incredible APIs and capabilities i never wouldve have sought out otherwise. congrats on 20 years!
100% - mixed with intellisense you get to explore a lib a lot more than just pasting a whole section
this is great. i think extensions that detect generated music, speech, video, or text will become really important. im curious how light and performant these detection models can get. maybe a single extension could handle multiple media types.
one concern (speaking as someone who doesnt know what these internal pipelines look like) is that suno/udio could tweak their model weights just enough to change the fingerprint, making a detector obsolete with each new release (or even more simple - maybe just apply post processing? id imagine a small reverb could diffuse the content enough to make the fingerprint difficult to detect). that turns it into a cat‑and‑mouse game. if its cheaper for them to mutate models/tweak post processing than for others to train new detectors, they could spin up a new fingerprint every day.
dont forget blinkies!
was this written by ai?
Audio Processing | Basic effects, limited voices | Complex DSP, many simultaneous sources | 150-400%
what? the web audio api is incredibly robust and performant. there is certainly no cap on voices besides memory constraints. writing your own audio engine in wasm would certainly not provide any advantage.
also all the use of tables + odd/broken styling on the page are red flags