HN user

che_shr_cat

708 karma
Posts144
Comments10
View on HN
arxiv.org 2mo ago

Universal Transformers Need Memory: Depth-State Trade-Offs in Adaptive Recursive

che_shr_cat
1pts0
cloud-heavy-industries.com 4mo ago

Game: Storm Cloud Simulator / Pure HTML5

che_shr_cat
1pts0
biorxiviq.substack.com 4mo ago

The Selfish Ribosome

che_shr_cat
1pts0
arxiviq.substack.com 7mo ago

Embedded Universal Predictive Intelligence: a coherent framework for multi-agent

che_shr_cat
1pts0
gonzoml.substack.com 7mo ago

NeurIPS 2025 Best Papers in Comics: From Artificial Hivemind to 1000-Layer RL

che_shr_cat
3pts0
gonzoml.substack.com 8mo ago

Visualizing Research: How I Use Gemini 3.0 to Turn Papers into Comics

che_shr_cat
1pts0
arxiviq.substack.com 8mo ago

Arc Is a Vision Problem

che_shr_cat
1pts0
arxiviq.substack.com 8mo ago

AlphaResearch: Accelerating New Algorithm Discovery with Language Models

che_shr_cat
1pts0
arxiviq.substack.com 8mo ago

LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics

che_shr_cat
2pts1
arxiviq.substack.com 8mo ago

Nested Learning: The Illusion of Deep Learning Architectures

che_shr_cat
6pts0
arxiviq.substack.com 8mo ago

Context Engineering 2.0: The Context of Context Engineering

che_shr_cat
2pts0
arxiviq.substack.com 8mo ago

A Practitioner's Guide to Kolmogorov-Arnold Networks

che_shr_cat
1pts0
arxiviq.substack.com 8mo ago

Kimi Linear: An Expressive, Efficient Attention Architecture

che_shr_cat
1pts0
arxiviq.substack.com 8mo ago

The Principles of Diffusion Models (470-pages)

che_shr_cat
2pts0
arxiviq.substack.com 9mo ago

CaT Replaces CoT-SC / Compute as Teacher: Turning Inference Compute Into

che_shr_cat
2pts0
gonzoml.substack.com 9mo ago

Tiny Recursive Model (TRM) vs. Hierarchical Reasoning Model (HRM)

che_shr_cat
2pts0
arxiviq.substack.com 9mo ago

Barbarians at the Gate: How AI Is Upending Systems Research

che_shr_cat
2pts0
arxiviq.substack.com 9mo ago

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

che_shr_cat
2pts0
arxiviq.substack.com 9mo ago

Autoreview: The Dragon Hatchling – The Missing Link Between the Transformer and

che_shr_cat
2pts0
gonzoml.substack.com 9mo ago

Stochastic Activations

che_shr_cat
2pts0
arxiviq.substack.com 10mo ago

LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures

che_shr_cat
1pts0
arxiviq.substack.com 10mo ago

Review: SpikingBrain Technical Spiking Brain-Inspired Large Models

che_shr_cat
2pts0
arxiviq.substack.com 10mo ago

K2-Think: A Parameter-Efficient Reasoning System

che_shr_cat
2pts1
arxiviq.substack.com 10mo ago

Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of AI

che_shr_cat
1pts0
arxiviq.substack.com 10mo ago

Fantastic Pretraining Optimizers and Where to Find Them

che_shr_cat
2pts0
arxiviq.substack.com 10mo ago

Solving the compute crisis with physics-based ASICs

che_shr_cat
5pts0
arxiviq.substack.com 10mo ago

Critiques of World Models

che_shr_cat
2pts0
arxiviq.substack.com 11mo ago

DeepConf: Scaling LLM reasoning with confidence, not just compute

che_shr_cat
98pts35
gonzoml.substack.com 11mo ago

V-JEPA 2: Scaling V-JEPA

che_shr_cat
2pts0
arxiviq.substack.com 11mo ago

Speed Always Wins: A Survey on Efficient Architectures for Large Language Models

che_shr_cat
3pts0

I'm the author of this blog. That's correct, the texts are generated and then validated manually by me.

I also do manual reviews (https://gonzoml.substack.com/), but there are many more papers for which I don't have time to write a review. So I created a multi-agentic system to help me, and I'm constantly iterating to improve it. And I like the result. It was also validated by the paper authors a couple of times, they agree the reviews are correct. So, if you see something is definitely wrong, please let me know.

Regarding myself, I became at least x10 more productive in reading papers and understanding what's happening. Hope, it will also help some of you.

That means the models still confabulate and make other errors, and for many cases it's a problem, so there should be solutions to control model output quality