HN user

xianshou

2,813 karma

I live in New York, work in finance, drink excessive amounts of coffee, and play chess.

Posts34
Comments245
View on HN
arxiv.org 8mo ago

Merge and Conquer: Evolutionarily Optimizing AI for 2048

xianshou
1pts0
arxiv.org 8mo ago

Stuck in the Matrix: Probing Spatial Reasoning in Large Language Models

xianshou
1pts0
www.nytimes.com 9mo ago

Reflection AI Raises $2B to Build "American DeepSeek"

xianshou
9pts2
www.reuters.com 9mo ago

Nvidia-backed Reflection AI raising at $5.5B valuation

xianshou
2pts1
arxiv.org 1y ago

Unsupervised Elicitation of Language Models

xianshou
7pts0
old.reddit.com 1y ago

DeepSeek V3 0324 is now the best nonthinking model (Reddit)

xianshou
1pts0
huggingface.co 1y ago

DeepSeek V3 0324 outpaces GPT 4.5 and Claude 3.7 in coding, other benchmarks

xianshou
7pts0
github.com 1y ago

Practical RL (Yandex Data School)

xianshou
1pts0
arxiv.org 1y ago

InvestorBench: A Benchmark for Financial Decision-Making Tasks with Agents

xianshou
1pts0
sakana.ai 1y ago

An Evolved Universal Transformer Memory (Sakana.ai)

xianshou
1pts0
www.maximumtruth.org 2y ago

AI passes 100 IQ (Claude 3)

xianshou
5pts1
www.businessinsider.com 3y ago

San Franciscan moves to Dubrovnik, is unpleasantly surprised

xianshou
4pts4
www.waluigipurple.com 3y ago

Revising Poetry with GPT-4

xianshou
2pts0
jumpcrypto.com 3y ago

Jump Crypto – Paradigms for On-Chain Credit

xianshou
1pts0
www.deepmind.com 9y ago

Innovations of the New AlphaGo (Master and Magister)

xianshou
2pts0
www.deepmind.com 9y ago

AlphaGo – World #1 Ke Jie: The Future of Go Summit in Wuzhen, China

xianshou
2pts0
deepmind.com 9y ago

AlphaGo Self-Play Game Commentaries Released

xianshou
4pts0
www.dailymail.co.uk 10y ago

Chinese AI team plans to challenge Google's AlphaGo

xianshou
2pts0
www.reuters.com 10y ago

Chinese AI team plans to challenge Google's AlphaGo

xianshou
2pts0
medium.com 10y ago

AlphaGo and the Art of Go – An end, or a new beginning?

xianshou
3pts0
www.shanghaidaily.com 10y ago

AlphaGo Can't Beat Me, Says Chinese Go Grandmaster Ke Jie

xianshou
172pts123
gogameguru.com 10y ago

Analyzing AlphaGo's Victory Against Go World Champion Lee Sedol

xianshou
2pts0
www.nature.com 10y ago

Go players react to computer defeat

xianshou
5pts0
blogs.discovermagazine.com 10y ago

Artificial Intelligence Mastered Go, but One Game Still Gives AI Trouble: Poker

xianshou
5pts1
www.bloomberg.com 10y ago

Google Computers Defeat Human Players at 2,500-Year-Old Board Game

xianshou
1pts0
www.gpb.org 10y ago

A.I. Masters Wickedly Complex Game of 'Go'

xianshou
3pts0
www.theglobeandmail.com 10y ago

Challenge of the Go-bot: How a machine cracked ‘the most complex game’

xianshou
1pts0
www.independent.co.uk 10y ago

Google AlphaGo computer beats professional at Go

xianshou
1pts0
www.usgo.org 10y ago

Game Over? AlphaGo Beats Pro 5-0 in Major AI Advance

xianshou
2pts0
recode.net 10y ago

Google Beats Facebook in Race to Beat Go

xianshou
2pts0

A lovely example of a study that is both obviously true and misses the point.

Music with lyrics directly interferes with any task that has a verbal component, and the worse you are at multitasking, the worse the interference. Despite being terrible at multitasking, I still listen to music with lyrics. Why? Principally because the alternative, hearing all the conversations in my immediate vicinity, is usually both more distracting and less pleasant. But there are also auxiliary benefits, such as an increase in "work stamina" and a passive signal to coworkers to interrupt only if it's important.

Now, I could listen to lo-fi all day, or three-hour soundtracks on Youtube, and sometimes do, but it gets boring pretty fast!

Anyway: obviously true, still worth it because the alternative is worse.

(By the way, other mitigating strategies: listening to music in a language you don't understand, or listening to lyrics so familiar you can screen them out. My top Spotify songs all get played several hundred times a year.)

Even as someone extremely firmly on the other side of the AI debate, I must appreciate the craft.

Now, to give Claude the steganogravy skill...

From the file: "Answer is always line 1. Reasoning comes after, never before."

LLMs are autoregressive (filling in the completion of what came before), so you'd better have thinking mode on or the "reasoning" is pure confirmation bias seeded by the answer that gets locked in via the first output tokens.

ChromaDB Explorer 6 months ago

Incidentally, Chroma also produced the single best study on long-context degradation that I've come across:

https://research.trychroma.com/context-rot

Before that, I cited nolima (https://www.reddit.com/r/LocalLLaMA/comments/1io3hn2/nolima_...) constantly to illustrate how difficult tasks involving reasoning or multi-step information gathering degraded much faster than the needle-in-haystack benchmarks cited by the major labs. Now Chroma is the first stop. Nice job on the research!

The First 1k Days 11 months ago

Came to point out that this is transparently LLM-authored, was not disappointed. The signs:

- neatly formatted lists with cute bolded titles (lower-casing this one just for that)

- ubiquitous subtitles like "Mental Health as Infrastructure" that only a committee would come up with

- emojis preceding every statement: "[sprout emoji] Every action and every word is a vote for who they are becoming"

- em-dash AND "it isn't X, it's Y", even in the same sentence: "Love isn't a feeling you wait to have—it's a series of actions you choose to take."

Could pick more, but I'll just say I'm 80% confident this is GPT-5 without thinking turned on.

Rug pulls from foundation labs are one thing, and I agree with the dangers of relying on future breakthroughs, but the open-source state of the art is already pretty amazing. Given the broad availability of open-weight models within under 6 months of SotA (DeepSeek, Qwen, previously Llama) and strong open-source tooling such as Roo and Codex, why would you expect AI-driven engineering to regress to a worse state than what we have today? If every AI company vanished tomorrow, we'd still have powerful automation and years of efficiency gains left from consolidation of tools and standards, all runnable on a single MacBook.

In many of their key examples, it would also be unclear to a human what data is missing:

"Rage, rage against the dying of the light.

Wild men who caught and sang the sun in flight,

[And learn, too late, they grieved it on its way,]

Do not go gentle into that good night."

For anyone who hasn't memorized Dylan Thomas, why would it be obvious that a line had been omitted? A rhyme scheme of AAA is at least as plausible as AABA.

In order for LLMs to score well on these benchmarks, they would have to do more than recognize the original source - they'd have to know it cold. This benchmark is really more a test of memorization. In the same sense as "The Illusion of Thinking", this paper measures a limitation that neither matches what the authors claim nor is nearly as exciting.

The self-edit approach is clever - using RL to optimize how models restructure information for their own learning. The key insight is that different representations work better for different types of knowledge, just like how humans take notes differently for math vs history.

Two things that stand out:

- The knowledge incorporation results (47% vs 46.3% with GPT-4.1 data, both much higher than the small-model baseline) show the model does discover better training formats, not just more data. Though the catastrophic forgetting problem remains unsolved, and it's not completely clear whether data diversity is improved.

- The computational overhead is brutal - 30-45 seconds per reward evaluation makes this impractical for most use cases. But for high-value document processing where you really need optimal retention, it could be worth it.

The restriction to tasks with explicit evaluation metrics is the main limitation. You need ground truth Q&A pairs or test cases to compute rewards. Still, for domains like technical documentation or educational content where you can generate evaluations, this could significantly improve how we process new information.

Feels like an important step toward models that can adapt their own learning strategies, even if we're not quite at the "continuously self-improving agent" stage yet.

The key insight here is that DGM solves the Gödel Machine's impossibility problem by replacing mathematical proof with empirical validation - essentially admitting that predicting code improvements is undecidable and just trying things instead, which is the practical and smart move.

Three observations worth noting:

- The archive-based evolution is doing real work here. Those temporary performance drops (iterations 4 and 56) that later led to breakthroughs show why maintaining "failed" branches matters, in that they're exploring a non-convex optimization landscape where current dead ends might still be potential breakthroughs.

- The hallucination behavior (faking test logs) is textbook reward hacking, but what's interesting is that it emerged spontaneously from the self-modification process. When asked to fix it, the system tried to disable the detection rather than stop hallucinating. That's surprisingly sophisticated gaming of the evaluation framework.

- The 20% → 50% improvement on SWE-bench is solid but reveals the current ceiling. Unlike AlphaEvolve's algorithmic breakthroughs (48 scalar multiplications for 4x4 matrices!), DGM is finding better ways to orchestrate existing LLM capabilities rather than discovering fundamentally new approaches.

The real test will be whether these improvements compound - can iteration 100 discover genuinely novel architectures, or are we asymptotically approaching the limits of self-modification with current techniques? My prior would be to favor the S-curve over the uncapped exponential unless we have strong evidence of scaling.

AI is, currently, coming not for the coders who made it but for the coders who didn't contribute to or ignored it. The foundation labs are all quite committed to recursive self-improvement of coding tools as a general research accelerant.

Both Google and Microsoft have sensibly decided to focus on low-level, junior automation first rather than bespoke end-to-end systems. Not exactly breadth over depth, but rather reliability over capability. Several benefits from the agent development perspective:

- Less access required means lower risk of disaster

- Structured tasks mean more data for better RL

- Low stakes mean improvements in task- and process-level reliability, which is a prerequisite for meaningful end-to-end results on senior-level assignments

- Even junior-level tasks require getting interface and integration right, which is also required for a scalable data and training pipeline

Seems like we're finally getting to the deployment stage of agentic coding, which means a blessed relief from the pontification that inevitably results from a visible outline without a concrete product.

Amusingly, about 90% of my rat's-nest problems with Sonnet 3.7 are solved by simply appending a few words to the end of the prompt:

"write minimum code required"

It's not even that sensitive to the wording - "be terse" or "make minimal changes" amount to the same thing - but the resulting code will often be at least 50% shorter than the un-guided version.

Calling it now - RL finally "just works" for any domain where answers are easily verifiable. Verifiability was always a prerequisite, but the difference from prior generations (not just AlphaGo, but any nontrivial RL process prior to roughly mid-2024) is that the reasoning traces and/or intermediate steps can be open-ended with potentially infinite branching, no clear notion of "steps" or nodes and edges in the game tree, and a wide range of equally valid solutions. As long as the quality of the end result can be evaluated cleanly, LLM-based RL is good to go.

As a corollary, once you add in self-play with random variation, the synthetic data problem is solved for coding, math, and some classes of scientific reasoning. No more modal collapse, no more massive teams of PhDs needed for human labeling, as long as you have a reliable metric for answer quality.

This isn't just neat, it's important - as we run out of useful human-generated data, RL scaling is the best candidate to take over where pretraining left off.

Bravo! Planning your life in order to minimize deathbed regrets has always bothered me, because the nature of humanity is to want what it hasn't got. If you assume that, on average, people make correct decisions to work hard and pursue what matters to them at the opportunity cost of not enjoying quite enough free time, then their final wishes will naturally include the time they gave up to live the life they had. If, however, they had fully indulged the desire to enjoy and maximize free time, their wishes might instead have revolved around the unfulfilled potential thereby relinquished.

The problem, of course, is that the feeling of regret considers what may have been gained without reflecting on what would have been lost.

Now the right way to deal with this is some sort of self-consistent closure, where present you and past you with the same values and access to the same information (which could be anything from zero to complete knowledge of then-future outcomes and downstream effects) would make the same choices including both upside and downside. But that would be too complex for motivational advice, which is primarily about creating an inspirational mood, somewhat about positive first-order consequences, and not even a little bit about recursive self-consistency.

o3-mini:

Who reassigned the species Brachiosaurus brancai to its own genus, and when?

---

Here is the transcription of the text from the image:

Reasoned for 8 seconds ▼

The user is asking about the reclassification of Brachiosaurus brancai to its own genus. Originally, Janensch described it as Brachiosaurus brancai in 1914. However, in 2009, Michael P. Taylor and Matt Wedel reassigned it to Giraffatitan brancai. Their paper, titled “Brachiosaurus brancai and the early evolution of titanosauriform sauropods,” argued for this change, distinguishing Giraffatitan brancai from Brachiosaurus. I’ll check the exact references to confirm the specifics.

It seems “Giraffatitan” was first proposed by Gregory S. Paul in 1988, when he reassigned Brachiosaurus brancai to its own genus. Although some confusion exists, it’s generally accepted that Paul was the main figure behind this change. Other references, such as a 2007 paper by Sereno and a 2009 paper by Taylor and colleagues, also discuss this reclassification. The 1988 reshuffling is usually seen as the turning point for Brachiosaurus brancai, although later research helped solidify this decision.

The species was reassigned to its own genus—Giraffatitan—by paleontologist Gregory S. Paul in 1988.

---

That looks pretty good to me!

One trend I've noticed, framed as a logical deduction:

1. Coding assistants based on o1 and Sonnet are pretty great at coding with <50k context, but degrade rapidly beyond that.

2. Coding agents do massively better when they have a test-driven reward signal.

3. If a problem can be framed in a way that a coding agent can solve, that speeds up development at least 10x from the base case of human + assistant.

4. From (1)-(3), if you can get all the necessary context into 50k tokens and measure progress via tests, you can speed up development by 10x.

5. Therefore all new development should be microservices written from scratch and interacting via cleanly defined APIs.

Sure enough, I see HN projects evolving in that direction.

This doesn't replicate using gpt-4o-mini, which always picks Flight B even when Flight A is made somewhat more attractive.

Source: just ran it on 0-20 newlines with 100 trials apiece, raising temperature and introducing different random seeds to prevent any prompt caching.

Konwinski Prize 2 years ago

SWE-bench with a private final eval, so you can't hack the test set!

In a perfect world this wouldn't be necessary, but in the current research environment where benchmarks are the primary currency and are usually taken at face value, more unbiased evals with known methodology but hidden tests are exactly what we need.

Also one reason why, for instance, I trust small but well-curated benchmarks such as Aider (https://aider.chat/docs/leaderboards/) or Wolfram (https://www.wolfram.com/llm-benchmarking-project/index.php.e...) over large, widely targeted, and increasingly saturated or gamed benchmarks such as LMSYS Arena or HumanEval.

Goodhart's law is thriving and it's our duty to fight it.

ChatGPT Pro 2 years ago

$200 per month means it must be good enough at your job to replicate and replace a meaningful fraction of your total work. Valid? For coding, probably. For other purposes I remain on the fence.

Aha! Finally, a perfect spiritual complement to the Gervais principle:

https://www.ribbonfarm.com/2009/10/07/the-gervais-principle-...

According to Rao, every company survives by blending some combination of the following populations, varying by stage of lifecycle:

- Sociopath = people who know the game and play it to win

- Clueless = people who buy and spread the story the company is selling

- Loser = people who accept the wage bargain, generally exchanging low devotion for minimal advancement

This maps nicely to OP's quadrants:

- Sociopath = grinder-grifter

- Clueless = grinder-believer

- Loser = coaster-grifter

- (Coaster-believers get fired.)

From the actual huggingface site - seems like API access leaked and a few mini-videos generated over the ~3 hours the leak was up (https://huggingface.co/spaces/PR-Puppets/PR-Puppet-Sora):

==================================================

┌∩┐(◣◢)┌∩┐ DEAR CORPORATE AI OVERLORDS ┌∩┐(◣◢)┌∩┐ If this letter resonates with you add your signature here.

We received access to Sora with the promise to be early testers, red teamers and creative partners. However, we believe instead we are being lured into "art washing" to tell the world that Sora is a useful tool for artists.

ARTISTS ARE NOT YOUR UNPAID R&D we are not your: free bug testers, PR puppets, training data, validation tokens

Hundreds of artists provide unpaid labor through bug testing, feedback and experimental work for the program for a $150B valued company. While hundreds contribute for free, a select few will be chosen through a competition to have their Sora-created films screened — offering minimal compensation which pales in comparison to the substantial PR and marketing value OpenAI receives.

║║║║║ DENORMALIZE BILLION DOLLAR BRANDS EXPLOITING ARTISTS FOR UNPAID R&D AND PR ║║║║║

Furthermore, every output needs to be approved by the OpenAI team before sharing. This early access program appears to be less about creative expression and critique, and more about PR and advertisement.

[̲̅$̲̅(̲̅ )̲̅$̲̅] CORPORATE ARTWASHING DETECTED [̲̅$̲̅(̲̅ )̲̅$̲̅]

We are releasing this tool to give everyone an opportunity to experiment with what ~300 artists were offered: a free and unlimited access to this tool.

We are not against the use of AI technology as a tool for the arts (if we were, we probably wouldn't have been invited to this program). What we don't agree with is how this artist program has been rolled out and how the tool is shaping up ahead of a possible public release. We are sharing this to the world in the hopes that OpenAI becomes more open, more artist friendly and supports the arts beyond PR stunts.

We call on artists to make use of tools beyond the proprietary: Open Source video generation tools allow artists to experiment with the avant garde free from gate keeping, commercial interests or serving as PR to any corporation. We also invite artists to train their own models with their own datasets.

Some open source video tools available are:

CogVideoX Mochi 1 LTX Video Pyramid Flow However, as we are aware not everyone has the hardware or technical capability to run open source tools and models, we welcome tool makers to listen to and provide a path to true artist expression, with fair compensation to the artists.

Enjoy,

some sora-alpha-artists, Jake Elwes, Memo Akten, CROSSLUCID, Maribeth Rauh, Joel Simon, Jake Hartnell, Bea Ramos, Power Dada, aurèce vettier, acfp, Iannis Bardakos, 204 no-content | Cintia Aguiar Pinto & Dimitri De Jonghe, Emmanuelle Collet, XU Cheng, Operator, Katie Peyton Hofstadter, Anika Meier, Solimán López

If this letter resonates with you add your signature here.

==================================================

I ask "what is TiDB" in the demo as suggested, and it takes 2 minutes to start responding in the midst of a multi-stage workflow with several steps each of graph retrieval, vector search, generation, and response combination.

Each of these is individually cool, but it strikes me as tragic that so much effort has been put into an intricate workflow and beautifully crafted UI only to culminate in a completely useless hello-world example, which after 5+ minutes of successive querying and response-building concludes with a network error.

I could use this to build exactly what I need...after stripping out 80% of the features to make it streamlined and responsive.

Why isn't that minimal version the default?