I'm waiting for an agent evaluated on a vending benchmark to start hacking into banks and wiring more money to its account so it can do better business.
HN user
isusmelj
Deep learning enthusiast interested in solving real world challenges.
I think this is the most satisfying video I’ve seen of a robot doing laundry. Not sure if it’s the camera angle, the 3× speed, or the music.
Not sure if it’s just me, but with Anthropic, every new feature has metered usage and “unlimited spending” (aka no limit) enabled by default for our team org. So if I activate something (Claude Code, Claude Tag) and don’t actively go to the usage page to set a spending limit, there is no limit. I’m not surprised Anthropic makes a lot of money. Most people in a typical org probably don’t even know how to check usage. Now, using Claude via Slack will just escalate that even further. And the only models I can use are Opus 4.7 and 4.8 for Claude Tag. If someone knows how to change the defaults, please let me know. With OpenAI, it’s the opposite. Everything is part of the plan, and by default there is no extra spending.
No note about the specific GPU they use. One might speculate. B200? H200? H100?
Is it just me, or does it feel like everyone now uses AI to write any kind of blog?
These parts here somehow trigger me:
- Enter TorchTPU. As an engineering team, our mandate was to build a stack that leads with usability, portability, and excellent performance.
- Engineering the TorchTPU Stack: The Technical Reality
- Eager First: Flexibility Without Compromise
- The breakthrough, however, is our fused eager mode.
- The Road Ahead: 2026 and Beyond
I have mixed feelings about this. On one hand, we all seem to be using the same tools and converging to the same style. On the other hand, if we all use the same models with the same system prompts, we might lose a lot of creativity and diversity in online content.
Yes, Marble (from World Labs) feels like it's generating Gaussian Splats or similar. I guess it's more compatible and easier to use for 3d asset generation and reusing in other software. Very exciting times ahead!
Is it just me or is the page barely readable? Lots of text is light grey on white background. I might have "dark" mode on on Chrome + MacOS.
I just wanted to check whether there is any information about the pricing. Is it the same as Qwen Max? Also, I noticed on the pricing page of Alibaba Cloud that the models are significantly cheaper within mainland China. Does anyone know why? https://www.alibabacloud.com/help/en/model-studio/models?spm...
I think we are just very close to the peak of a typical Gartner hype cycle around LLMs. They are useful but overhyped. There will be more posts about fuckups that happen because people run things on autopilot and cannot keep up with reviewing AI generated code.
Do not get me wrong. I use AI all day to speed things up. But I believe that there is only a small group, maybe 5 percent or less, that actually knows how to use AI properly (I'd count myself not yet in that 5%), which I see as potentially dangerous. The other issue I see is inexperienced software engineers writing software. Although I see this as a great value add and productivity boost for prototyping, I am afraid of the “I do not know much about coding but can also make PRs to our codebase” mentality.
For those of you that run things on autopilot, how do you keep code quality under control? And how do you handle refactoring? I am really curious, because one option now is also to just YOLO your LLMs to write code based on the maturity of the product. You can refactor an app or parts of it pretty fast again with LLMs. While tech debt accumulates faster, we also have the opportunity to rebuild faster.
Are there any benchmarks? I didn’t find any. It would be the first model update without proof that it’s better.
Very proud as a Swiss that Soumith has a .ch domain!
Is the price here correct? https://openrouter.ai/moonshotai/kimi-k2-thinking Would be $0,60 for input and $2,50 for 1 million output tokens. If the model is really that good it's 4x cheaper than comparable models. It's hosted at a loss or the others have a huge margin? I might miss something here. Would love some expert opinion :)
FYI: the non thinking variant has the same price.
I can only agree with your experience in Europe. I do not get how they do that, but Tesla Superchargers are more reliable. The occupancy information works better, they are easier to use, and they almost always offer a more competitive price. I often see other chargers that are 50 to 100 percent more expensive and only very rarely see offers that are within 10 to 50 percent.
What strikes me is that this difference can make EVs more expensive per kilometer if you only compare energy cost with fuel cost.
Here is the math with numbers. Tesla chargers in Switzerland and Germany are usually at most CHF 0.50 or EUR 0.60 per kilowatt hour at the more expensive locations, along highways for example. They offer fast charging of 150 kW or more. Alternative providers often start at around CHF 0.75 for 50 kW or CHF 1.00 for more than 250 kW fast charging. If your electric car consumes 20 kWh (Model 3 is at around 15 I think) per 100 km you end up with costs of CHF 10.00, CHF 15.00, or CHF 20.00 per 100 km at CHF 0.50, CHF 0.75, or CHF 1.00 per kilowatt hour. If you drive a petrol car that uses 8 l per 100 km and the cost per liter is CHF 1.70 you pay CHF 13.60 per 100 km.
Are there any news about power consumption? I didn’t even see a tdp or so mentioned.
Demand > Supply?
Thanks for clarifying! I wish you all the best luck!
I hope they do well. AFAIK they’re training or finetuning an older LLaMA model, so performance might lag behind SOTA. But what really matters is that ETH and EPFL get hands-on experience training at scale. From what I’ve heard, the new AI cluster still has teething problems. A lot of people underestimate how tough it is to train models at this scale, especially on your own infra.
Disclaimer: I’m Swiss and studied at ETH. We’ve got the brainpower, but not much large-scale training experience yet. And IMHO, a lot of the “magic” in LLMs is infrastructure-driven.
Is there anything like this also supporting other GPUs? Thinking of Apple Silicon or embedded ones in phones etc.
As someone in Europe, I sometimes wonder what’s worse: letting US companies use my data to target ads, or handing it to Chinese companies where I have no clue what’s being done with it. With one I at least get an open source model. The other is a big black box.
You're right, UncleEntity, thanks for highlighting that. My phrasing could have been clearer. AGPL does allow various uses, including commercial, provided its terms are met.
Our intention with LightlyTrain (AGPL/Commercial license option) is to offer a streamlined, production-ready pretraining engine. This contrasts with our other library, LightlySSL (github.com/lightly-ai/lightly), which is MIT-licensed and geared towards researchers needing flexible building blocks.
We found many companies wanted a simpler "it just works" solution for pretraining, which is why LightlyTrain exists with its specific licensing options tailored for commercial teams alongside the AGPL.
Thanks again for the clarification!
Hi Sonnigeszeug, great that you're looking into LightlyTrain!
We designed LightlyTrain specifically for production teams who need a robust, easy-to-use pretraining solution without getting lost in research papers. It builds on learnings from our MIT-licensed research framework, LightlySSL (github.com/lightly-ai/lightly), but is tailored for scalability and ease of integration.
For commercial use where the AGPL terms might not fit your needs, we offer straightforward commercial licenses for LightlyTrain. Happy to chat more if that's relevant for you!
Thanks for the kind words, joelio182! Glad you see the value in making SSL more practical for real-world domain shift issues.
As liopeer mentioned, we have results for medical (DeepLesion) and agriculture (DeepWeeds) in the blog post. We haven't published specific benchmarks on satellite or industrial inspection data yet, but those are definitely the kinds of niche domains where pretraining on specific unlabeled data should yield significant benefits. We're keen to explore more areas like these.
Our goal is exactly what you pointed out - bridging the gap between SSL research and practical application where labels are scarce. Appreciate the encouragement!
Hi HN, I’m Igor, co-founder of Lightly AI (https://www.lightly.ai/).
We just released LightlyTrain, a new open-source Python package (AGPL-3.0, free for research and educational purpose) for self-supervised pretraining of computer vision models: https://github.com/lightly-ai/lightly-train
Standard vision models pretrained on generic datasets like ImageNet or COCO often underperform on specific domains (e.g., medical, agriculture, autonomous driving). Fine-tuning helps, but performance is limited, and getting enough labeled data is expensive and slow.
LightlyTrain uses self-supervised learning (SSL) to pretrain models directly on your own unlabeled images or videos. This adapts the model to your specific visual domain before fine-tuning, leading to significantly better performance with less labeled data.
Key Features:
- No Labels Needed: Pretrain using your existing unlabeled image data.
- Better Performance: Consistently outperforms training from scratch and ImageNet-pretrained weights, especially in low-data regimes and domain-specific tasks (benchmarks in README/blog). We see gains across detection, classification, and segmentation.
- Domain Adaptation: Tailor models to your specific industry (manufacturing, healthcare, retail, etc.).
- Supports Popular Models: Works out-of-the-box with YOLO (v5-v12), RT-DETR, ResNet, ViTs, etc., integrating with frameworks like Ultralytics, TIMM, Torchvision.
- Easy to Use & Scalable: Simple pip install, minimal code to start, scales to millions of images, runs fully on-premise (single/multi-GPU). We built this because while SSL research is mature, making it easily accessible and effective for industry computer vision teams was hard. LightlyTrain aims to bridge that gap.
We’ve benchmarked it on COCO, BDD100K (driving), DeepLesion (medical), and DeepWeeds (agriculture), showing strong improvements over baselines (details in the repo/blog post linked below). For example, on COCO with only 10% labels, LightlyTrain pretraining boosted YOLOv8-s mAP by +14% over ImageNet weights and +34% over no pretraining.
- GitHub Repo: https://github.com/lightly-ai/lightly-train
- Docs: https://docs.lightly.ai/train
- Detailed Blog Post/Benchmarks: https://www.lightly.ai/blog/introducing-lightly-train
- Quick Demo Video: https://youtu.be/5Lmry1k_cA8
We’re here to answer any questions! Happy to discuss the tech, benchmarks, or use cases. Commercial licenses are also available for businesses needing different terms.
I’ve been playing around with SIMD since uni lectures about 10 years ago. Back then I started with OpenMP, then moved to x86 intrinsics with AVX. Lately I’ve been exploring portable SIMD for a side project where I’m (re)writing a Numpy-like library in Rust, mostly sticking to the standard library. Portable SIMD has been super helpful so far.
I’m on an M-series MacBook now but still want to target x86 as well, and without portable SIMD that would’ve been a headache.
If anyone’s curious, the project is here: https://github.com/IgorSusmelj/rustynum. It's just a learning exercise for learning Rust, but I’m having a lot of fun with it.
Does anyone know why GPT4 has knowledge cutoff December 2023 and all the other models (newer ones like 4o, O1, O3) seem to have knowledge cutoff October 2023? https://platform.openai.com/docs/models#o3-mini
I understand that keeping the same data and curating it might be beneficial. But it sounds odd to roll back in time with the knowledge cutoff. AFAIK, the only event that happened around that time was the start of the Gaza conflict.
I think the results show that just in general the compute is not used well. That the CPU took 8.4ms and GPU took 3.2ms shows a very small gap. I'd expect more like 10x - 20x difference here. I'd assume that the onnxruntime might be the issue. I think some hardware vendors just release the compute units without shipping proper support yet. Let's see how fast that will change.
Also, people often mistake the reason for an NPU is "speed". That's not correct. The whole point of the NPU is rather to focus on low power consumption. To focus on speed you'd need to get rid of the memory bottleneck. Then you end up designing your own ASIC with it's own memory. The NPUs we see in most devices are part of the SoC around the CPU to offload AI computations. It would be interesting to run this benchmark in a infinite loop for the three devices (CPU, NPU, GPU) and measure power consumption. I'd expect the NPU to be lowest and also best in terms of "ops/watt"
I think that is similar to what Yann LeCun outlined: https://bdtechtalks.com/2022/03/07/yann-lecun-ai-self-superv...
Don't forget that this is 24k H100. They are getting 10x the compute: https://www.cnbc.com/2024/01/18/mark-zuckerberg-indicates-me...
So gpt-4 level 8B models running on phones and notebooks seems feasible within the next 5 years. I imaging having (voice) assistans running locally. Crazy how fast we progress.
Is there somewhere an overview of the progress we made on the software side for training and inference of LLMs? It feels like we squeezed 10-100x more out of the hardware since llama appeared. This crazy progress will probably saturate though as we reach theoretical limits, no?
This is really cool.
Song: https://app.suno.ai/song/83680b6f-db37-44de-adf9-3f7fff6b79d...
Prompt: A 90s hip-hop song with a male singer with a deep voice singing about how AI models are creating new songs after being trained on all the data of artists. Talks about AI models are stealing the show.
Lyrics:
[Verse] Step back, my friend, 'cause the future's here AI models spittin' rhymes, oh so clear (oh so clear) Trained on data from all the greats Now they're droppin' beats that dominate (dominate)
[Verse 2] Don't need no ghostwriters or melody makers These AI models are the true risk-takers (oh yeah) Analyzin' every flow, every precise word Stealing the show, that's just absurd (it's absurd)
[Chorus] AI takin' over, breakin' the mold Stolen styles, but they're icy cold (they're icy cold) The game's been changed, no human control The rise of the AI flow, takin' its toll (takin' its toll)