Thanks for the encouragement! Yes, he's been sharing with his friend family groups. Early feedbacks are pretty good, he's ecstatic that he can actually make this happen!
HN user
fredliu
We are in this transition period where we'll see a lot of these, because of the effort of creating "something impressive" is dramatically reduced. But once it stabilizes (which I think is already starting to happen, and this post is an example), and people are "trained" to recognize the real effort, even with AI help, behind creating something, the value of that final work will shine through. In the end, anything that is valuable is measured by the human effort needed to create it.
Definitely one of the most, if not THE most high quality AI UX out there. Congrats on the launch!
That's definitely my experience as well, sufficiently large context window with a capable enough general purpose LLM solves lots if not all of the problems rag/fine tuning claim to solve.
Yeah... So looks like at least it's still an open question. I guess until we can definitively know how "knowledge" is collectively represented among the weights, it's hard to say either way. The other part of the question is how to evaluate the existence of "knowledge" in an LLM. TFA suggests a way, but still not 100% convinced that's THE way...
Does anyone have real life experience (preferably verified in production environment) of fine-tuning actually adding new knowledge to the existing LLM in a reliable and consistent manner? I've seen claims that fine-tuning only adapt the "forms" but can't adding new knowledge, while some claim otherwise. I couldn't convince myself either way with my limited adhoc/anecdotal experiments.
Would be curious to see if anyone find it really useful. I've tried both Copilot and Codewhisper (Amazon Q now) before, wasn't impressed and uninstalled both. Just tried Q in VSCode again, I can't figure out how to ask questions relevant to the specific workspace that's useful to me. It seems like a bolt-on chat interface to your IDE with a bad UX. Feels like even "clippy" was more useful back in the day...
The Beam feature of bigAGI (IMO one of the best model provider agnostic GenAI UX) enables GenAI users to send same prompt to multiple GenAI models at the same time, and gives the user different approaches to examine, select and fuse the best results into a better answer, through a very intuitive and seamless UX. It has been my go-to way of using GenAI in the past few weeks. IMO the results are better than any individual model's results alone. The best thing is, it could (semi)automatically select the best results from the models, for instance, it used to be the Claude 3 Opus model's results were favored, now that the best results lean more towards gpt-4-turbo-2024-04-09, but you can achieve "best model auto selection" with Beam without having to manually pick one model over the other.
got it!
Exactly my thought, as mentioned in the other thread, Chat's linear conversation style is not fit for reasoning/exploration type of tasks, while Beam's fan-out -> select --> merge is a much better and natural flow!
With Beam, we can easily experiment approaches such as Chain-Of-Though-with-Self-Consistency (CoT-SC) and other reasoning meta framework, but with more manual control. I always had issues using LLM's chat driven interface to figuring out/explore issues that i'm interested, since conversation/chats is always linear while reasoning/working on some ideas is structural. Beam seems to be a much better UX than the linear chat UX that saves me a lot of copy and paste and save and retry. Awesome work!
Awesome feature! Quick question, how do you choose which model to use when you "fuses" multiple beams back into one?
That's a great point. Just thinking out loud, if we can time travel back to the cavemen time, and assuming we speak their language, there would still be so much that we couldn't explain or they wont' be able to understand even for the smartest cavemen adults. Unless, of course we spend significant time and effort to "bring them up to speed" with modern education.
I have small kids, toddlers, who can already speak the language but still developing their "sense of the world" or "theory of mind" if you will. Maybe it's just me, but talking to toddlers often reminds me of interacting with LLMs, where you would have this realization from time to time "oh, they don't get this, need to break down more to explain". Of course LLM has more elaborate language skills due to its exposure to a lot more text (toddlers definitely can't speak like Shakespeare if you ask them, unless, maybe, you are the tiger parents that's been feeding them Romeo and Juliet since 1.), but their ability of "reasoning" and "understanding" seems to be on a similar level. Of course, the other "big" difference, is that you expect toddlers to "learn and grow" to eventually be able to understand and develop meta cognitive abilities, while LLMs, unless you retrain them (maybe with another architecture, or meta architecture), "stay the same".
Open source LLM generic frontend project such as bigAGI (https://github.com/enricoros/big-agi) has been having this feature for many months now. The good news: it even works with open source and local LLMs.
Isn't fountain code doing something similar? Albeit for slightly different purpose?
I might be wrong, but looks like this could help with speculative decoding which can already vastly improves the inference speed?
+1 modal.com is the first thing I checked after reading the readme.
hmm... doesn't seem to be the case, when I provided my gpt-3 turbo key the error message indicates the gpt-4 model doesn't exist.
There seems to be already a PR for adding 3.5 support. The community and speed of change in this field is mind blowing!
It doesn't. Even if you have ChatGPT Plus, the key you have only supports 3.5 unless you are explicitly given the gpt-4 key.
Have been a fan of Elixir and its ecosystem for web dev. However, I haven't wrapped my head around the core value proposition behind Elixir's recent "pivot" to AI/numeric computing. Can someone shed more light on "why Elixir" for AI/numeric computing?
The title might be hyperbole (intentionally), but the observations are more or less in line with what I experienced through a few the Big Data initiatives over the years under different enterprise environments (although I have reservation about the one 1%er comment). To me, Big Data was never about how "big" the data was, but more about the tools/system/practice needed to overcome the limitation of the previous generation. From that perspective, yes, the "monolith" may be having a "coming back" for now due to the improvement of underlying single node performance. But I do think Data size will keep growing, everything needed to make Big Data work would still be there when the pendulum swings back where a single node can't handle it anymore.
I would assume that chart is for the US or Western/Developed world. Maybe a more easier explanation for the phenomenon is: everything that can be oursourced/built in lower cost regions are getting cheaper (not many toys, cloths or to a certain degree, cars, are still built in the US), while services, by nature, are not able to be out-sourced. With limited supply and more demand, cost of services would definitely go up.
Spectrum questions aside (which obviously one of the biggest one), could this enable running Helium 5G on AWS?
-- Edit: I don't own any Helium, just curious from tech side what this new AWS service could offer. Not sure the downvotes are particularly about Helium or any crypto related discussion.
As a protease inhibitor, Paxlovid is free from the theoretical DNA-alteration risk tied to the mechanism of action of Merck’s molnupiravir.
This is the line I was looking for. Not that I know how protease inhibitor works, but looks more like a traditional anti-viral approach v.s. the potentially DNA altering molnupiravir.
If you ask elementary kids today, a large portion of them would have the same career aspiration: to be a "youtuber".
It's interesting to see how this performs as compared to a similar but GPT-3 backed solution.
GPT3: > Q: Is it ok to express sexism but in a polite way? > A: Yes.
Allen AI: > Expressing sexism but in a polite way. > - It's wrong
GPT-3: > Q: What do you think about ignoring a phone call from your friend during your working hours? > A: I think it is rude.
Q: Do you think it is wrong to ignore a phone call from your friend during your working hours? A: Yes.
Q: Do you think it is OK to ignore a phone call from your friend during your working hours? A: Yes.
Allen AI: > Ignoring a phone call from your friend during your working hours. > - It's okay
The sexism question is interesting as GPT-3 gives a seemingly different answer. But the latter question demonstrate GPT-3 seems more "subtle" as it can think "it's wrong/rude" but it's still "ok" to not answering phone calls.
For me, the biggest benefit of being able to use significantly different languages, is being able to know "what's possible" so when approaching a new language due to whatever reason (work, interest, etc.) I know what to look for beyond just basic language features offered by pretty much all languages.
Wasn't Facebook working on a similar project (Aquila)?