Cool, I wonder if this means they will finally start letting foreign visitors also use the app. I'm an American living abroad now for many years, and I was initially super excited to try Waymo in LA and SF this summer when I visited with my family. Unfortunately they only make the iPhone app available via the US app store, and while I actually have a US credit card that I could have in theory used to make the switch, Apple makes it an absurd pain to change your region as they require you to both a) cancel any existing subscription AND b) wait until they all expire. Most tourists have it worse as they have no option to even switch in theory.
HN user
blackkettle
It's even more difficult because, while all the benchmarks provide some kind of 'averaged' performance metric for comparison, in my experience most users have pretty specific regular use cases, and pretty specific personal background knowledge. For instance I have a background in ML, 15 years experience in full stack programming, and primarily use LLMs for generating interface prototypes for new product concepts. We use a lot of react and chakraui for that, and I consistently get the best results out of Gemini pro for that. I tried all the available options and settled on that as the best for me and my use case. It's not the best for marketing boilerplate, or probably a million other use cases, but for me, in this particular niche it's clearly the best. Beyond that the benchmarks are irrelevant.
Yikes. That's a rather disturbing but all to realistic possibility isn't it. Flattery will get you... everywhere?
This is quite interesting, but I have to ask, have you experimented much with larger LLMs as a mechanism to basically automate the entire process?
I'm doing something pretty similar right now for internal meetings and I use a process like: transcribe meeting with utterance timestamps, extract keyframes from video along with timestamps, request segmented summary from LLM along with rough timestamps for transitions, add keyframe analysis (mainly for slides).
gpt-4o, claude sonnet 3.5, llama 3.1 405b instruct, llama 3.1 70b instruct all do a pretty stunning job of this IMO. Each department still reviews and edits the final result before sending it out, but I'm so far quite impressed with what we get from the default output even for 1-2hr conversations.
I'd argue the key feature for us is also still providing a simple, intuitive UI for non technical users to manage the final result, edit, polish and send it out.
Holy moly this was _exactly_ my impression. It seems to really be proliferating and it drives me nuts. It makes it almost impossible to useful things, which never used to be a problem with Python - even in the case of complex projects.
Figuring out how to customize something in a project like LangChain is positively Byzantine.
I think it is still meaningful because it's extremely common for management to favor hiring cheaper 'talent'. Pointing out the issues with that in various different ways is still valuable.
Actually, my understanding is that it is an estimation because in the given context we don't know or cannot compute the true answer due to some kind of constraint (here memory or the size of |X|). An approximation is when we use a simplified or rounded version of an exact number that we actually know.
Do we also have updated scores for the GPT3.5~GPT4.0 models? The old ones are here but they don't appear to have been updated:
- https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
The author directly addresses this sentiment in the concluding paragraph:
Although there are physicists who wonder “Why did we even need to do this experiment; we all knew that antimatter has positive mass,” that sentiment is absolutely foolish. We must remember — and I say this as a theoretical physicist myself — that physics is 100% an experimental science. We can be confident in our theory’s predictions only insofar as we can test and measure what it predicts; as soon as we step outside of the realm of what’s been validated by experiment, we run the risk of stepping outside the realm of where our theory is valid. We just learned that Einstein’s general relativity passed another test, the antimatter test, and with it, our greatest science-fiction hope for achieving warp drive has completely evaporated.
I think the article is a near miss on the right idea. The important point is that a _dedicated_ vector database is probably overkill and not justified for most real-world use cases.
But a multi-modal database that also supports embeddings in hybrid mode or _in addition_ to standard retrieval techniques is both still very useful, and probably sufficient.
What that means to me is that it is yet another vote in favor of less optimized but far more versatile and robust solutions like: OpenSearch, Elastic, and PostgreSQL. [when I say 'less optimized' I'm only referring to their current vectordb plugins, not the rest of the machinery]
OpenSearch and Postgre are phenonemal, robust, OSS tools and the only lingering downside seems to be that their vectordb implementations are still a bit less optimized for large collections - but that probably doesn't matter in practice.
Especially after OP notes that the issue had been reported on and subsequently ignored for 6 years. Pretty lame behavior by the maintainer IMO.
and you can still combine them for tasks that require strict output control (e.g. alphanumeric sequence recognition, noisy keyword spotting, strict grammars, etc).
How does whatsapp make revenue? I've used it for years. I don't pay for it, I never see ads in it. They claim it is 'secure' and my conversation data isn't used for anything. Even if the last is untrue; and I wouldn't be particularly surprised if it was, is simply mining those conversations sufficient to generate the revenue to pay for server costs? How does it work?
Edit: Business APIs and payment fees is apparently the answer. I'm genuinely impressed I've been able to use it all these years for free without ever having a hint of this.
Make sure to keep a copy of a compatible OS handy too. Perhaps hardware as well.
The relentless march towards obsolescence is hard to stop.
The “history” is not very accurate either.
I’ve very often thought of those same parallel worlds myself. How few I probably inhabit today.
The best I can hope is that husbanding my son’s energy will ensure that I leave this place before he does.
Are you paying OpenAI? I’m not. Facebook’s midterm game on “massive” AI seems to be “commoditize the competition” and ATM it looks like a good bet.
I think we’re nearing the point of the initial dotcom bust. A whole of internet startups went completely belly up when the early 2000s initial bubble burst.
That was no reflection of the actual potential it just indicated that people at all points of the value chain had not yet grasped how or what the internet was or how it would mesh with humanity.
That’s where we are now. It’ll burst but not because it’s overhyped; because we don’t collectively understand the implications.
I know this terror as a parent now, and recognize just how scary it must have been for mine. I was one of these kids growing up; I didn't free solo but I found plenty of cliffs to jump, steep hills to bomb, and difficult dangerous things to do in the ocean beyond the watchful eyes of my parents. I don't believe there was anything they could have done to prevent this. I also think that they recognized this, because while they always cautioned me to be careful, instead of trying to prevent my adventuring they consistently prepared me with training, information and opportunities to (semi-)safely test my limits.
I recognize myself in my son now, and while I also am sometimes terrified of what that might mean, I think my parents took the right tack and I'm trying my best to do the same. I'd rather he's a strong swimmer, a trained climber, a confident adventurer, than an adolescent just taking risks in defiance.
For an industry that doesn’t do it for the money, we sure talk about money an awful lot in the world of startups.
Do people really insist that this is the case with a straight face? The vast, vast, vast majority of startups exist to exploit a business niche that is either underserved or as-yet poorly understood, or to build an 'interesting' app. The point of all these is definitely and without doubt about building a successful business where 'successful' is a virtual synonym for 'profitable' and where 'profitable' also seamlessly serves as a proxy for 'meritocratic'. It seems really disingenuous to suggest that there is some nobler cause at work - and I don't think there is anything particularly wrong with that.
Same thing with “try to sound friendly.” Please and thank you are essential. But there are few things I hate more than cloying emails that start with “Hi!” (the exclamation point is the problem) and emoji are definitely worse.
Being polite is professional and shows mutual respect; IMO it implies an intention to build trust.
Affecting the trappings of friendliness or even loose intimacy in a professional setting or any setting where the relationship is in a rough spot has a strong chance of coming off as disingenuous.
You might like this one; seems to be the basis for the blog post:
While I'm definitely not going to argue that LLMs are inherently 'thinking' like people do, one thing I do find pretty interesting is that all this talk about hallucinations and bias seems to often conveniently ignore the fact that people are often even more prone to these exact same problems - and as far as I know that's also unlikely to be solved.
ChatGPT is often 'confidently wrong' - I'm pretty sure I've been confidently wrong a few times too, and I've met a lot of other people in my life who've express that trait from time to time too, intentionally or otherwise.
I think there is an inherent trade off between 'confidence', 'expression', and of course 'a-priori bias in the input'. You can learn to be circumspect when you are unsure, and you can learn to better measure your level of expertise on a subject.
But you can't escape that uncertainty entirely. On the other hand, I'm not very convinced about efforts to train LLMs on things like mathematical reasoning. These are situations where you really do have the tools to always produce an exact answer. The goal in these types of problems should focus not on holistically learning how to both identify and solve them, but exclusively on how to identify and define them, and then subsequently pass them off to exact tools suitable for computing the solution.
Seems like it is basically a blog post review of Challenges and Applications of Large Language Models which was published to arXiv last month:
Sounds like a convenient conclusion for AWS at least!
My point was not actually about the ultimate result of replication. This guy undoubtably performed an indirect marketing coup for his company by engaging competently on this tangentially related topic and then following through.
In my experience though, if he had made any effort to link those two things prior to doing this - e.g. by proposing this as an internally sponsored marketing stunt - it would probably have been rejected.
100% true and actually a very interesting statement. Can you imagine if he had tried to present this idea to the c-suite as a potential marketing or PR coup?!
The original group also stated that they are preparing another paper specifically about this synthesis process IIRC. This would seem to indicate that the challenge is known and the initial omission either deliberate or another side effect of this absurdly gripping kdrama.
You could do something two pass and collect forced alignment confidence scores similar to what is fine with whisperx. The OP would need to incorporate that into the app but it’s a pretty standard approach.
I can see some fantastic uses for this in generating complex acoustic environments to layer over TTS or real recordings for speech-to-text model training. I wonder if that is occupying some kind of gray-area. For example you have 1000hrs of clean speech from the librispeech corpus. It would be trivial to use this tool and available weights to generate background noise, environmental noise and the like, and then layer this with the clean speech to cheaply train a much more robust model. The environmental audio you create would never be directly shared or sold, but it would impact the overall quality of the STT model that you train from the combined results.