HN user

tim_sw

12,831 karma

Yet another Tim

Posts1,386
Comments92
View on HN
blog.ezyang.com 6mo ago

The Gap Between a Helpful Assistant and a Senior Engineer

tim_sw
2pts0
www.aleksagordic.com 10mo ago

VLLM: Anatomy of a High-Throughput LLM Inference System

tim_sw
4pts1
arxiv.org 1y ago

DeepSeek: Inference-Time Scaling for Generalist Reward Modeling

tim_sw
163pts35
nicholas.carlini.com 1y ago

AI forecasting retrospective: you're (probably) over-confident

tim_sw
3pts0
scaling-reasoning.chrisbarber.co 1y ago

Will Scaling Reasoning Models Like O3 and R1 Unlock Superhuman Reasoning?

tim_sw
4pts0
www.lesswrong.com 1y ago

The Online Sports Gambling Experiment Has Failed

tim_sw
8pts1
github.com 1y ago

ByteDance's Recommendation System

tim_sw
64pts55
lexfridman.com 1y ago

Dario Amodei: Anthropic CEO on Claude, AGI and the Future of AI and Humanity

tim_sw
1pts0
waymo.com 1y ago

Waymo Incorporates Gemini for End to End Autonomous Driving Model

tim_sw
1pts0
waymo.com 1y ago

Waymo's Research on an End-to-End Multimodal Model for Autonomous Driving

tim_sw
5pts1
aws.amazon.com 1y ago

Automated reasoning often makes systems more efficient and easier to maintain

tim_sw
1pts0
justjake.substack.com 1y ago

Destructive Tightening

tim_sw
2pts1
www.bloomberg.com 1y ago

How the US Lost the Solar Power Race to China

tim_sw
1pts0
nlp.stanford.edu 1y ago

Instruction Following Without Instruction Tuning

tim_sw
2pts0
medium.com 1y ago

A Cynic's Guide to Fintech

tim_sw
1pts0
giansegato.com 1y ago

The Dawn of a New Startup Era

tim_sw
2pts1
www.lambrospetrou.com 1y ago

Durable Objects (DO) – Unlimited single-threaded servers spread across the world

tim_sw
3pts0
github.com 1y ago

Karpathy/Nano-Llama31

tim_sw
74pts1
sh4dy.com 1y ago

Writing a system call tracer using eBPF

tim_sw
99pts2
huyenchip.com 1y ago

Building a Generative AI Platform

tim_sw
1pts0
nostarch.com 2y ago

Writing a C Compiler – Build a Real Programming Language from Scratch

tim_sw
3pts1
dioxus.notion.site 2y ago

Dioxus Labs and "High-Level Rust"

tim_sw
4pts0
www.cnbc.com 2y ago

OpenAI ex-employees worry about company's control over their shares

tim_sw
3pts0
www.fabricatedknowledge.com 2y ago

The Internet as You Know It Is Dying

tim_sw
7pts4
www.oreilly.com 2y ago

What We Learned from a Year of Building with LLMs (Part II)

tim_sw
25pts4
i-admin.cetico.org 2y ago

From Ground Zero to Production: Go's Journey at Google

tim_sw
1pts0
bloomberry.com 2y ago

I scraped all of ChatGPT's Enterprise customers – here's what I learned

tim_sw
2pts1
stack.convex.dev 2y ago

How Convex Works

tim_sw
96pts28
venge.net 2y ago

Rust programming language (a.k.a. "Project Servo") [pdf]

tim_sw
3pts0
www.runtime.news 2y ago

The U.S. government is losing trust in Microsoft

tim_sw
39pts14

My 2 cents - don’t rely on these frameworks and just do it yourself (or pick libraries like Instructor over these frameworks)

I think both have the wrong abstractions for people to build more complex workflows and use cases beyond demos.

This is more common than reported. I’ve lost a semi expensive present, and have heard anecdotal stories of people losing handbags, jewelry, higher end clothing/shoes.

It seems to be more prevalent in the US/TSA than other first world countries.

this is not a duplicate - it's not the announcement post, it's a deep dive into the paper from the POV of Hugging Face's RL lead.

The devil is in the details. Training large LLMs requires a lot of custom infra (handling GPUs going down, efficiently pushing data to keep the accelators busy, deciding on which mechanism of parallelizing model training is better - data vs model parallelism or both, tuning hyperparams of optimizers which can be different for larger batch sizes, etc)

Mosaic is one of the better providers for this. AWS is nowhere near ready at this current point in time, it is pretty much a "dumb" infra provider in large LLM training at this point. (Of course they won't be standing still and will prob acquire that capability one way or another)

We find that responses from existing generative search engines are fluent and appear informative, but frequently contain unsupported statements and inaccurate citations: on average, a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence. We believe that these results are concerningly low for systems that may serve as a primary tool for information-seeking users, especially given their facade of trustworthiness. We hope that our results further motivate the development of trustworthy generative search engines and help researchers and users better understand the shortcomings of existing commercial systems.

Pretty high model flop utilization:

—————

MaxText is a high performance, arbitrarily scalable, open-source, simple, easily forkable, well-tested, batteries included LLM written in pure Python/Jax and targeting Google Cloud TPUs. MaxText typically achieves 55% to 60% model-flop utilization and scales from single host to very large clusters while staying simple and "optimization-free" thanks to the power of Jax and the XLA compiler.