HN user

zone411

4,258 karma

https://twitter.com/LechMazur

10 LLM benchmarks: https://github.com/lechmazur/

https://www.linkedin.com/in/lech-mazur-69b70493/

Advameg (City-data.com) founder and CEO. AI startup founder.

Author: AI melody songwriting assistant https://melodies.ai

Author: Accurate COVID-19 county-by-county neural net case prediction model based on most data.

Posts74
Comments817
View on HN
www.proofatlas.ai 1d ago

Natural-Density Almost-Bounded Collatz Orbits in Logarithmic Time (AI, Lean)

zone411
2pts0
github.com 3mo ago

LLM Position Bias Benchmark: Swapped-Order Pairwise Judging

zone411
1pts0
github.com 3mo ago

Show HN: Buyout Game Benchmark: Multi-Agent Bargaining, Transfers, and Takeovers

zone411
6pts0
github.com 3mo ago

LLM Persuasion Benchmark: Multi-Turn Persuasion Between Models

zone411
9pts0
github.com 4mo ago

Show HN: LLM Debate Benchmark

zone411
9pts3
github.com 4mo ago

Show HN: LLM Sycophancy Benchmark: Opposite-Narrator Contradictions

zone411
3pts0
github.com 10mo ago

Show HN: LLM Round‑Trip Translation Benchmark

zone411
6pts0
github.com 10mo ago

Show HN: LLM Creative Story‑Writing Benchmark V3

zone411
8pts0
github.com 10mo ago

Show HN: Mapping LLM Style and Range in Flash Fiction

zone411
7pts0
github.com 11mo ago

Pact: Head-to-head negotiation benchmark for LLMs

zone411
6pts0
github.com 1y ago

Show HN: Bazaar – a new LLM benchmark for economic reasoning under uncertainty

zone411
8pts1
www.quantamagazine.org 1y ago

AI Comes Up with Physics Experiments. But They Work

zone411
4pts0
github.com 1y ago

Emergent Price-Fixing by LLM Auction Agents

zone411
7pts0
github.com 1y ago

Public Goods Game Benchmark: Contribute and Punish, a Multi-Agent Benchmark

zone411
7pts0
github.com 1y ago

Elimination Game: Multi-Agent LLM Social Reasoning, Strategy, and Deception

zone411
5pts0
arxiv.org 1y ago

SWE-Lancer: a benchmark of freelance software engineering tasks from Upwork

zone411
111pts74
github.com 1y ago

LLM Hallucination Benchmark: R1, o1, o3-mini, Gemini 2.0 Flash Think Exp 01-21

zone411
17pts3
github.com 1y ago

Multi-Agent Step Race Benchmark: LLM Collaboration and Deception Under Pressure

zone411
7pts1
github.com 1y ago

Show HN: LLM Thematic Generalization Benchmark

zone411
6pts0
github.com 1y ago

Show HN: LLM Creative Story-Writing Benchmark

zone411
5pts0
github.com 1y ago

Show HN: LLM Divergent Thinking Creativity Benchmark

zone411
8pts0
github.com 1y ago

Show HN: LLM Deceptiveness and Gullibility Benchmark

zone411
7pts1
github.com 1y ago

LLM Confabulation (Hallucination) Leaderboard

zone411
6pts0
twitter.com 1y ago

O1-preview and o1-mini results on NYT Connections

zone411
2pts1
twitter.com 2y ago

Grok is an AI modeled after the Hitchhiker’s Guide to the Galaxy

zone411
213pts226
parrotchess.com 2y ago

Can you beat a stochastic parrot? ParrotChess.com

zone411
3pts4
labs.google.com 2y ago

Generative AI while browsing in Chrome

zone411
3pts0
www.safe.ai 3y ago

Statement on AI Risk

zone411
341pts921
www.businessinsider.com 3y ago

Google tells staff it plans to limit publishing AI research

zone411
63pts28
www.servethehome.com 3y ago

4th Gen Intel Xeon Scalable Sapphire Rapids Leaps Forward

zone411
2pts1

I actually tried using GPT-5.5 Pro on this problem recently. It thought it was making progress on one path, but it made so many mistakes that it didn't feel worth it pushing further. It'll be interesting to check whether it's the same route. I got partial results (proved in Lean) that improve on the best-known results for four Erdős problems with GPT-5.5 Pro

I built this benchmark this month: https://github.com/lechmazur/sycophancy. There are large differences between LLMs. There are large differences between LLMs. For example, Mistral Large 3 and GPT-4.1 will initially agree with the narrator, while Gemini will disagree. I swap sides, so this is not about possible viewpoint bias in the LLMs. But another benchmark shows that Gemini will then change its view very easily in a multi-turn conversation while Kimi K2.5 or Grok won't: https://github.com/lechmazur/persuasion.

GPT-5.4 5 months ago

Results from my Extended NYT Connections benchmark:

GPT-5.4 extra high scores 94.0 (GPT-5.2 extra high scored 88.6).

GPT-5.4 medium scores 92.0 (GPT-5.2 medium scored 71.4).

GPT-5.4 no reasoning scores 32.8 (GPT-5.2 no reasoning scored 28.1).

Claude Sonnet 4.6 5 months ago

They're improved compared to 4.5 on my Extended NYT Connections benchmark (https://github.com/lechmazur/nyt-connections/).

Sonnet 4.6 Thinking 16K scores 57.6 on the Extended NYT Connections Benchmark. Sonnet 4.5 Thinking 16K scored 49.3.

Sonnet 4.6 No Reasoning scores 55.2. Sonnet 4.5 No Reasoning scored 47.4.

GPT-5.2 7 months ago

I've benchmarked it on the Extended NYT Connections benchmark (https://github.com/lechmazur/nyt-connections/):

The high-reasoning version of GPT-5.2 improves on GPT-5.1: 69.9 → 77.9.

The medium-reasoning version also improves: 62.7 → 72.1.

The no-reasoning version also improves: 22.1 → 27.5.

Gemini 3 Pro and Grok 4.1 Fast Reasoning still score higher.

You got many answers already, but a couple more points:

Poker doesn't require lying or table talk. Bluffing is rule-legal strategic deception expressed through betting. More like a feint in sports than cheating.

If "sitting at a table following rules" is the issue, that's true of most games. And formats vary: many are short and cash games are leave-anytime.

The exact questions are almost certainly not in the training data, since extra words are added to each puzzle, and I don't publish these along with the original words (though there's a slight chance they used my previous API requests for training).

To guard against potential training data contamination, I separately calculate the score using only the newest 100 puzzles. Grok 4 still leads.