HN user

GaggiX

4,818 karma
Posts141
Comments1,821
View on HN
ilialarchenko.com 2h ago

Prizewinning Solution of the LeHome Challenge

GaggiX
1pts0
arxiv.org 20d ago

Explaining Attention with Program Synthesis

GaggiX
2pts0
arxiv.org 1mo ago

FlashMemory-DeepSeek-V4

GaggiX
1pts1
ideogram.ai 1mo ago

Ideogram 4.0 Technical Details: Open model at the forefront of design

GaggiX
3pts0
blog.chrislewis.au 1mo ago

Snowboard Kids 2 is 100% Decompiled

GaggiX
284pts109
huggingface.co 2mo ago

Nvidia: Nemotron Labs Diffusion 14B

GaggiX
3pts0
www.cerebras.ai 2mo ago

Cerebras Brings Trillion Parameter Inference to Enterprises with Kimi K2.6

GaggiX
2pts0
arxiv.org 4mo ago

Generalized Discrete Diffusion from Snapshots

GaggiX
1pts0
github.com 4mo ago

Attention Residuals

GaggiX
241pts34
www.anthropic.com 4mo ago

Eval awareness in Claude Opus 4.6's BrowseComp performance

GaggiX
3pts0
matharena.ai 5mo ago

MathArena: Evaluating LLMs on uncontaminated math questions

GaggiX
2pts0
static.stepfun.com 5mo ago

Step 3.5 Flash

GaggiX
2pts0
bfl.ai 6mo ago

FLUX.2 [Klein]: Towards Interactive Visual Intelligence

GaggiX
227pts58
showlab.github.io 7mo ago

OmniPSD: Layered PSD Generation with Diffusion Transformer

GaggiX
1pts0
natex-ldm.github.io 8mo ago

NaTex: Seamless Texture Generation as Latent Color Diffusion

GaggiX
2pts0
arxiv.org 8mo ago

Back to Basics: Let Denoising Generative Models Denoise

GaggiX
3pts0
developers.googleblog.com 9mo ago

Veo 3.1 and new creative capabilities in the Gemini API

GaggiX
4pts0
github.com 9mo ago

Linus Learns Analog Circuits

GaggiX
171pts42
unigen-x.github.io 10mo ago

UnifoLM-WMA-0: A World-Model-Action (WMA) Framework Under UnifoLM Family

GaggiX
2pts0
z.ai 12mo ago

GLM-4.5: Reasoning, Coding, and Agentic Abililties

GaggiX
247pts134
www.androidauthority.com 1y ago

A retro gaming YouTuber faces possible jail time for reviewing gaming handhelds

GaggiX
12pts0
arxiv.org 1y ago

From Bytes to Ideas: Language Modeling with Autoregressive U-Nets

GaggiX
2pts0
lizhihao6.github.io 1y ago

Sparse Representation and Construction for High-Resolution 3D Shapes Modeling

GaggiX
2pts0
lllyasviel.github.io 1y ago

Packing Input Frame Context in Next-Frame Prediction Models for Video Generation

GaggiX
270pts27
fiction.live 1y ago

Fiction.LiveBench: The First Long Context Benchmark for Writers

GaggiX
2pts0
jianhongbai.github.io 1y ago

ReCamMaster: Camera-Controlled Generative Rendering from a Single Video

GaggiX
1pts0
horwitz.ai 1y ago

Charting and Navigating Hugging Face's Model Atlas

GaggiX
1pts0
arxiv.org 1y ago

Block Diffusion: Interpolating between autoregressive and diffusion models

GaggiX
156pts32
www.igenius.ai 1y ago

iGenius Releases Colosseum 355B

GaggiX
2pts1
aim-uofa.github.io 1y ago

Framer: Interactive Frame Interpolation

GaggiX
1pts0

I really like Gemini 2.5 Flash Lite because it's a dirt cheap model that support every input modalities.

At least now MiMo v2.5 exists and can be used as another dirt cheap multimodal model.

I would be more interested about terrorists organization like Al-Shabaab that at least control many towns.

Does Boko Haram and ISWAP even control a single town or they just control a few villages in Lake Chad and in the Sambisa forest?

Also reading the report they seem quite clueless.

Kagi Magic 1 month ago

This is just an Ad, I thought it was a new product from Kagi.

Well with a standard autoregressive model you can generate for example 256 tokens at once if you have 256 users, with this approach you can generate 256 tokens for a single user but you need several forward steps.

So the diffusion process takes more GFLOPs, if you have enough users you can already balance memory and compute.

Not to be confused with Flash Attention.

What's novel here is the extremely small KV cache memory usage per long context windows, like 0.77GB with 512K, a 90% memory usage reduction compare to the already really small KV cache memory usage of Deepseek V4 Flash.

AI is slowing down 1 month ago

I’ll take a few f bombs and the truth.

Don't want to ruin it but go read some old posts from the author about AI, the tone is the same and he is very much wrong.

That's technically encoding

Isn't that just projecting the patches into the d_model size vectors that the models takes?

I am assuming that involves of quantization

12B model in 16GB seems very reasonable to me, int8 is top quality for running models.

Dav2d 2 months ago

Is opus being used for the audio or it's not the solution for extreme lossy compression?

Dav2d 2 months ago

I love this, hope to see a AV2 version at 8MB

Dav2d 2 months ago

I would love to see comparisons with AV1 on very low bitrates.

The giant Umarell in the background is a nice piece of furniture.

Edit: I noticed later it was in Milan, I guess it makes perfect sense.

Gemini 3.5 Flash 2 months ago

If don't want to spend 1.5$/9$ for the lastest model then yes use a cheaper model, DeepSeek V4 Flash is 0.11$/0.22$ on OpenRouter and it's more capable than the most expensive model a year ago. Models have never been so cheap given their capabilities unless you want to follow the SOTA (where the hype is).