What do you think about Grok 4.5 in comparison to Muse Spark 1.1?
HN user
axiom92
https://madaan.github.io
This was integrated in gpt4 2 years ago:
https://www.reddit.com/r/ChatGPT/comments/14sqcg8/anyone_els...
And even before this work, there was "PAL: Program-aided Language Models" (https://arxiv.org/abs/2211.10435, https://reasonwithpal.com/).
Afaik PaLM (Google's OG big models) tried this trick, but it didn't work for them. I think it's because PaL used descriptive inline comments + meaningful variable names. Compare the following:
```python
# calculate the remaining apples
apples_left = apples_bought - apples_eaten
```
vs.
```python
x = y - z
```
We have ablations in https://arxiv.org/abs/2211.10435 showing that both are indeed useful (see "Crafting prompts for PAL").
You can do this at grok.com.
There is a "start thread" option below every conversation. You can also read the responses aloud (helpful if you want to do something async).
Yeah, that's the first formal reference I remember as well (although, BERT is probably the first thing NLP folks will think of after reading about diffusion).
I collected a few other text-diffusion early references here about 3 years ago: https://github.com/madaan/minimal-text-diffusion?tab=readme-....
From last neurips https://automix-llm.github.io/automix/
The demo was done live (as was everything else).
Looks pretty cool https://www.youtube.com/watch?v=mogSbMD6EcY
Although, it seems it's only going to cover the first book (which makes sense, given how difficult the other two would be to film). The real magic for me was in book 3. It was inspiring to see someone think so far out, so boldly.
Welcome to one of the most hated parts of the academia.
The joke is that he doesn't own any OpenAI shares.
[1] https://www.cnbc.com/2023/03/24/openai-ceo-sam-altman-didnt-...
Right, but no separate image encoder + half the size could be very helpful for many applications.
tasks that need deterministic outputs and the thing you need to create is already known statically
Wow, interesting. Do you have any example for this?
I've realized that LLMs are fairly good at string processing tasks that a really complex regex might also do, so I can see the point in those.
And basically all servers will have 8xA100
for those wondering: no this is not the norm. My lab at CMU doesn't own any A100s (we have A6000s).
The options seemed to be: If I went for it, I’d be penniless, and if I didn’t go for it, I’d be bitter. I’d be bitter going forward. Penniless certainly beats bitter. So I made the decision.
Kind of like industry -> PhD decision.
Some of our recent/relevant work: https://selfrefine.info/
For those curious about self-refining systems: https://selfrefine.info/ (our recent work).
cushman and code-davinci are similar for sure (same architecture). Perhaps that's what they meant.
Codex (code-davinci-002) is free (limited beta).
Actually, there is no way to be sure^. If you think about the costs + scale, it's likely to be cushman (code-cushman-001).
^ Unless you are from OpenAI, in which case I have more questions for you :)
codex (code-davinci-002) was free. This is going to be a huge deal for research groups.
We did some work in exploring why spelling out the rationale before the answer works so well!
Talk: https://madaan.github.io/res/presentations/TwoToTango.pdf
Sure, but we don't know if ChatGPT is based on the original GPT-3 architecture.
As they also mention, this is not the first time FTC has done this. Here is the earlier AI guidance from 4/2021: https://www.ftc.gov/business-guidance/blog/2021/04/aiming-tr....
Ummm not really? Any decent language model will produce sentences that look like legit instructions given this user's prompts.
It has been possible to generate impressive graphs from text since GPT-2. Though you need a few tricks to make it work.
Here's an example (my work): https://aclanthology.org/2021.naacl-main.67.pdf
TLDR of the input/output: https://madaan.github.io/res/tldr/graph_gen_tldr.jpg
Some work from AllenAI: https://proscript.allenai.org/
I see. I wonder if you've been phrasing this (tricky) question correctly.
For example, if you've been asking "I don't smell bad right?? I smell great, right!?" you're unlikely to get honest replies.
Reminds me of "The Mom Test" [1]. The book has a few tricks that can help with asking tough questions and receiving an honest feedback.
until your microbiome is able to handle all the waste / oil that your skin produces
Or maybe until you stop noticing the smell? Ask a friend perhaps.
Wow, that's a great negative sample for any reading class--uselessly abstruse and verbose.
Took me multiple turns even with ChatGPT to simplify it:
Using desire for control may seem manageable at first, but it leads to negative consequences and people try to justify it using false reasoning and fake authority to make it seem acceptable, even though it goes against rational thinking.