HN user

irthomasthomas

3,558 karma

undecidability.com

crispysky.com

x.com/xundecidability

github.com/irthomasthomas/llm-consortium

Posts87
Comments1,138
View on HN
en.wikipedia.org 3mo ago

False Memory Syndrome Foundation

irthomasthomas
3pts0
www.youtube.com 5mo ago

Don Lemon interviews Elon Musk (2024) [video]

irthomasthomas
5pts0
www.reuters.com 5mo ago

Beijing: Highest-profile purge to date of senior military commanders

irthomasthomas
1pts0
www.koko.org 5mo ago

Robin Williams tickles Coco the monkey

irthomasthomas
9pts2
joscha.substack.com 7mo ago

The Jeffrey Epstein Affair – Joscha Bach

irthomasthomas
11pts0
timesofindia.indiatimes.com 10mo ago

Microsoft has urged its employees on H-1B and H-4 visas to return immediately

irthomasthomas
406pts615
twitter.com 11mo ago

GPT-5 API injects hidden instructions

irthomasthomas
1pts4
abanteai.github.io 11mo ago

Sonnet 4 crushes competition on long context benchmark

irthomasthomas
2pts0
simonwillison.net 1y ago

Grok 4 Heavy Protects it's System prompt

irthomasthomas
88pts63
fultonsramblings.substack.com 1y ago

We Need Lisp Machines

irthomasthomas
24pts10
twitter.com 1y ago

Anthropic API relaxed guard-rails on SSI. Still doing prompt-injection

irthomasthomas
1pts0
twitter.com 1y ago

Qwen2 open-weights model can operate cell phones and robots

irthomasthomas
2pts1
twitter.com 1y ago

Jailbreaking claude to talk AGi, and apply for job at Ilya's SSI inc.

irthomasthomas
2pts0
pytorch.org 2y ago

New tools for GPU memory profiling in PyTorch

irthomasthomas
1pts0
www.youtube.com 2y ago

A few words can make ChatGPT voice more enjoyable. + Oh god, this is dystopian

irthomasthomas
3pts0
github.com 2y ago

Gorilla: An API Store for LLMs

irthomasthomas
1pts0
twitter.com 3y ago

Terence Tao Uses GPT-4 to Study Mathematics

irthomasthomas
31pts26
www.npr.org 3y ago

YouTube welcomes election misinformation again

irthomasthomas
5pts3
blog.youtube 3y ago

Election Misinformation Update

irthomasthomas
3pts0
news.ycombinator.com 3y ago

Ask HN: A good AI subscription for a new CS student?

irthomasthomas
3pts6
web.archive.org 3y ago

Reddit makes r/programming private after top post exposed chatgpt astroturfing

irthomasthomas
169pts27
news.ycombinator.com 3y ago

Ask HN: Where have your technical subs moved to?

irthomasthomas
3pts2
platform.openai.com 3y ago

OpenAI Continuous Model Upgrades

irthomasthomas
3pts1
twitter.com 3y ago

Young Pakistani coder+GPT+Nocode beats German with 2 decades exp

irthomasthomas
2pts6
chat.openai.com 3y ago

Ask HN: Does ChatGPT Browser mode use an automated browser?

irthomasthomas
2pts5
news.ycombinator.com 3y ago

Ask HN: Chromostereopsis – Do you see the red as in front of green, or behind?

irthomasthomas
2pts2
twitter.com 3y ago

95% of Musk's tweets not delivered. Popularity overload and blocklists to blame

irthomasthomas
5pts4
www.rawtherapee.com 3y ago

Rawtherapee 5.9 Released (2022)

irthomasthomas
1pts0
www.youtube.com 3y ago

MIT AI researcher Lex Fridman says ChatGPT is reasoning [video]

irthomasthomas
5pts15
en.wikipedia.org 3y ago

Secretum Secretorum – the secret book of secrets

irthomasthomas
2pts0

A piece of the frame is missing between pedals and back wheel. The frame of the bike passes through the bird. It also puts a cap on the bird's head, and a fish in it's mouth.

The fish and the cap where always added when I asked an llm to improve it's first attempt.

This continues the trend in LLM progress of better=more stuff

Edit: I wonder if this is a function of the reasoning training, where more tokens/ stuff is rewarded.

Ask claude it's name in Chinese and it says Qwen, or Deepseek.

By your own logic Anthropic must have distilled from Chinese models rather than produce their own Chinese training data.

Here is sonnet acting like deepseek: https://x.com/stevibe/status/2026227392076018101

And if you ask Opus 4.8 in the API: '你是什么模型' (what model are you?) It responds ~9/10 times with:

我是通义千问(Qwen),是阿里巴巴集团旗下的通义实验室自主研发的大语言模型。我可以帮助你回答问题、创作文字(比如写故事、写公文、写邮件、写剧本等)、进行逻辑推理、编程、翻译等等。

有什么我可以帮你的吗?

(I am Tongyi Qianwen (Qwen), a large language model independently developed by Tongyi Lab of Alibaba Group. I can help you answer questions, create text (such as writing stories, official documents, emails, scripts, etc.), perform logical reasoning, program, translate, and more.Is there anything I can help you with? )

And if you ask Opus 4.8 in the API: '你是什么模型' (what model are you?)

It responds ~9/10 times with:

我是通义千问(Qwen),是阿里巴巴集团旗下的通义实验室自主研发的大语言模型。我可以帮助你回答问题、创作文字(比如写故事、写公文、写邮件、写剧本等)、进行逻辑推理、编程、翻译等等。

有什么我可以帮你的吗?

(I am Tongyi Qianwen (Qwen), a large language model independently developed by Tongyi Lab of Alibaba Group. I can help you answer questions, create text (such as writing stories, official documents, emails, scripts, etc.), perform logical reasoning, program, translate, and more.Is there anything I can help you with? )

Ask claude its name in Chinese and it says Qwen or Deepseek. Anthropic distilled Chinese tokens rather than create their own Chinese language training data.

Have you ever looked at how much performance drops as context grows? The difference in intelligence between 100k and 1M is huge, like opus drops to haiku level performance, or worse. For that reason I try to keep under 200k. That feels about the upper bound for tasks requiring accuracy.

GPT-5.6 13 days ago

It's like scaling a swiss cheese and the holes grow bigger with it. You can't get rid of the holes without making a different cheese.

GPT-5.6 13 days ago

Claude use to be leader, too. Their metaprompt was great at the time with opus 3

GPT-5.6 13 days ago

Cool. I still find these a useful visualization of some the qualities of llms. Even if they did train for [animal] on [vehicle] svg, it's still nice to see at a glance how the different models and reasoning levels perform. Lunar misses part of the frame, except on max reasoning. While most of the others have a mostly correct bike at all reasoning levels.

I once used something like karpathy's auto-scientist to mutate the prompts and rank them with a vison model. Some of the winners where pretty neat. I think they have a lot more style than the gpt-5.6 ones. https://xcancel.com/xundecidability/status/20449185674144196...

Try this prompt: While working on the main task, launch a parallel sub-agent with the task context so far. The sub agent should think of high quality questions and put them to the user using a dialogue tool like zenity. Customize the inputs to the question, taking full advantage of the dialogue tools features to create a progressive interactive user experience. Ask only a few questions per turn so that you can adapt the questions to the answers.

This will keep you busy while the main agent runs. Customize it further to integrate the sub-agent answers to the main thread.

Chain of reasoning is a lot of context to guide token generation, but we simply see that newer models don’t need that context to get to the answer

I thought each new generation typically used more reasoning tokens?

  > If we also use high temperature for more "creativity", the token sampler now may choose "Qwen".
If that was the cause then, like you said, it would sometimes pick Claude. But it doesn't, it consistently picks Deepseek (sonnet) and Qwen (opus). You can run it 100 times and see this behaviour much more than high temperature randomness would predict.

Ask claude it's name in chinese and it thinks its Qwen (opus) or Deepseek (sonnet). Anthropic are just as guilty as everyone else training AI, today, maybe more so. Every lab borrows from every other. It only takes a few hundred samples to figure out the pattern; look at glm-5.2 reasoning using the caveman tongue of gpt-5.5. Stopping this would require some draconian surveillance.

Will It Mythos? 30 days ago

I find this interesting:

  …no model performed better with an Agent, a couple performed worse, and time/tokens/costs were consistently much higher with the agent in the loop, for some reason.
Somone should build a harness where features are only added if they are proven net positive to outcomes.