HN user

isusmelj

298 karma

Deep learning enthusiast interested in solving real world challenges.

Posts43
Comments57
View on HN
github.com 11mo ago

Show HN: Distill DINOv3 into your own model

isusmelj
1pts0
ai.meta.com 11mo ago

DINOV3: Self-supervised learning for vision at unprecedented scale

isusmelj
10pts0
github.com 1y ago

Show HN: LightlyTrain – Pretrain YOLO/ResNet on unlabeled data beats ImageNet

isusmelj
3pts0
www.lightly.ai 1y ago

Nvidia B200 vs. H100: Independent Training Performance Tests

isusmelj
6pts0
www.lightly.ai 1y ago

Clip and Friends: How Vision-Language Models Evolved

isusmelj
4pts0
www.lightly.ai 1y ago

Albumentations vs. PIL: 2x Speedup for Model Training Pipelines

isusmelj
4pts0
github.com 1y ago

Show HN: RustyNum – Fast Rust-Powered NumPy Alternative

isusmelj
1pts0
www.lightly.ai 2y ago

Image Retrieval using your fine-tuned DINO model

isusmelj
6pts0
www.lightly.ai 2y ago

Data Curation Demystified for Stable Video Diffusion

isusmelj
1pts0
github.com 2y ago

Show HN: Lightly Insights – open-source dataset analysis

isusmelj
2pts0
github.com 2y ago

Show HN: Lightly – A Python library for self-supervised learning on images

isusmelj
5pts0
github.com 2y ago

Show HN: Labelformat now supports all major vision labeling formats

isusmelj
1pts0
github.com 2y ago

Show HN: Labelformat – Python package for converting computer vision labels

isusmelj
11pts0
news.ycombinator.com 2y ago

Counting “Generative AI” Appearances on Google Cloud Next Page

isusmelj
1pts0
www.lightly.ai 3y ago

Apple M1 and M2 Performance for Training SSL Models

isusmelj
1pts0
www.lightly.ai 3y ago

Improve Your Large Language Models (LLMs) with Active Learning

isusmelj
2pts0
www.lightly.ai 3y ago

Boosting YOLOv8 Accuracy and Saving 77% in Labeling Costs with Active Learning

isusmelj
3pts0
github.com 3y ago

A Python library for self-supervised learning on images

isusmelj
1pts0
www.lightly.ai 3y ago

Self-supervised learning trends and what to expect in 2023

isusmelj
1pts0
www.lightly.ai 3y ago

Guide for active learning in computer vision

isusmelj
1pts0
twitter.com 3y ago

Can we paint using ChatGPT?

isusmelj
1pts0
cloud.google.com 3y ago

Google shared more specs around TPUv4 in their cloud

isusmelj
1pts0
news.ycombinator.com 3y ago

GitHub Down?

isusmelj
13pts10
twitter.com 3y ago

Example of deep learning models struggling with out-of-distribution data

isusmelj
4pts0
www.lightly.ai 4y ago

Self-Supervised Models Are More Robust and Fair

isusmelj
3pts0
www.lightly.ai 4y ago

Train Test Split in Deep Learning

isusmelj
6pts1
news.ycombinator.com 4y ago

Launch HN: Lightly (YC S21): Label only the data which improves your ML model

isusmelj
88pts25
www.lightly.ai 5y ago

Active Learning Using Detectron2

isusmelj
5pts1
lightly.ai 5y ago

The Advantage of Self-Supervised Learning

isusmelj
3pts1
lightly.ai 5y ago

Embedded Covid mask detection on an Arm Cortex-M7 processor using PyTorch

isusmelj
2pts0

I think this is the most satisfying video I’ve seen of a robot doing laundry. Not sure if it’s the camera angle, the 3× speed, or the music.

Claude Tag 29 days ago

Not sure if it’s just me, but with Anthropic, every new feature has metered usage and “unlimited spending” (aka no limit) enabled by default for our team org. So if I activate something (Claude Code, Claude Tag) and don’t actively go to the usage page to set a spending limit, there is no limit. I’m not surprised Anthropic makes a lot of money. Most people in a typical org probably don’t even know how to check usage. Now, using Claude via Slack will just escalate that even further. And the only models I can use are Opus 4.7 and 4.8 for Claude Tag. If someone knows how to change the defaults, please let me know. With OpenAI, it’s the opposite. Everything is part of the plan, and by default there is no extra spending.

Is it just me, or does it feel like everyone now uses AI to write any kind of blog?

These parts here somehow trigger me:

- Enter TorchTPU. As an engineering team, our mandate was to build a stack that leads with usability, portability, and excellent performance.

- Engineering the TorchTPU Stack: The Technical Reality

- Eager First: Flexibility Without Compromise

- The breakthrough, however, is our fused eager mode.

- The Road Ahead: 2026 and Beyond

I have mixed feelings about this. On one hand, we all seem to be using the same tools and converging to the same style. On the other hand, if we all use the same models with the same system prompts, we might lose a lot of creativity and diversity in online content.

I think we are just very close to the peak of a typical Gartner hype cycle around LLMs. They are useful but overhyped. There will be more posts about fuckups that happen because people run things on autopilot and cannot keep up with reviewing AI generated code.

Do not get me wrong. I use AI all day to speed things up. But I believe that there is only a small group, maybe 5 percent or less, that actually knows how to use AI properly (I'd count myself not yet in that 5%), which I see as potentially dangerous. The other issue I see is inexperienced software engineers writing software. Although I see this as a great value add and productivity boost for prototyping, I am afraid of the “I do not know much about coding but can also make PRs to our codebase” mentality.

For those of you that run things on autopilot, how do you keep code quality under control? And how do you handle refactoring? I am really curious, because one option now is also to just YOLO your LLMs to write code based on the maturity of the product. You can refactor an app or parts of it pretty fast again with LLMs. While tech debt accumulates faster, we also have the opportunity to rebuild faster.

I can only agree with your experience in Europe. I do not get how they do that, but Tesla Superchargers are more reliable. The occupancy information works better, they are easier to use, and they almost always offer a more competitive price. I often see other chargers that are 50 to 100 percent more expensive and only very rarely see offers that are within 10 to 50 percent.

What strikes me is that this difference can make EVs more expensive per kilometer if you only compare energy cost with fuel cost.

Here is the math with numbers. Tesla chargers in Switzerland and Germany are usually at most CHF 0.50 or EUR 0.60 per kilowatt hour at the more expensive locations, along highways for example. They offer fast charging of 150 kW or more. Alternative providers often start at around CHF 0.75 for 50 kW or CHF 1.00 for more than 250 kW fast charging. If your electric car consumes 20 kWh (Model 3 is at around 15 I think) per 100 km you end up with costs of CHF 10.00, CHF 15.00, or CHF 20.00 per 100 km at CHF 0.50, CHF 0.75, or CHF 1.00 per kilowatt hour. If you drive a petrol car that uses 8 l per 100 km and the cost per liter is CHF 1.70 you pay CHF 13.60 per 100 km.

Nvidia DGX Spark 11 months ago

Are there any news about power consumption? I didn’t even see a tdp or so mentioned.

I hope they do well. AFAIK they’re training or finetuning an older LLaMA model, so performance might lag behind SOTA. But what really matters is that ETH and EPFL get hands-on experience training at scale. From what I’ve heard, the new AI cluster still has teething problems. A lot of people underestimate how tough it is to train models at this scale, especially on your own infra.

Disclaimer: I’m Swiss and studied at ETH. We’ve got the brainpower, but not much large-scale training experience yet. And IMHO, a lot of the “magic” in LLMs is infrastructure-driven.

As someone in Europe, I sometimes wonder what’s worse: letting US companies use my data to target ads, or handing it to Chinese companies where I have no clue what’s being done with it. With one I at least get an open source model. The other is a big black box.

You're right, UncleEntity, thanks for highlighting that. My phrasing could have been clearer. AGPL does allow various uses, including commercial, provided its terms are met.

Our intention with LightlyTrain (AGPL/Commercial license option) is to offer a streamlined, production-ready pretraining engine. This contrasts with our other library, LightlySSL (github.com/lightly-ai/lightly), which is MIT-licensed and geared towards researchers needing flexible building blocks.

We found many companies wanted a simpler "it just works" solution for pretraining, which is why LightlyTrain exists with its specific licensing options tailored for commercial teams alongside the AGPL.

Thanks again for the clarification!

Hi Sonnigeszeug, great that you're looking into LightlyTrain!

We designed LightlyTrain specifically for production teams who need a robust, easy-to-use pretraining solution without getting lost in research papers. It builds on learnings from our MIT-licensed research framework, LightlySSL (github.com/lightly-ai/lightly), but is tailored for scalability and ease of integration.

For commercial use where the AGPL terms might not fit your needs, we offer straightforward commercial licenses for LightlyTrain. Happy to chat more if that's relevant for you!

Thanks for the kind words, joelio182! Glad you see the value in making SSL more practical for real-world domain shift issues.

As liopeer mentioned, we have results for medical (DeepLesion) and agriculture (DeepWeeds) in the blog post. We haven't published specific benchmarks on satellite or industrial inspection data yet, but those are definitely the kinds of niche domains where pretraining on specific unlabeled data should yield significant benefits. We're keen to explore more areas like these.

Our goal is exactly what you pointed out - bridging the gap between SSL research and practical application where labels are scarce. Appreciate the encouragement!

Hi HN, I’m Igor, co-founder of Lightly AI (https://www.lightly.ai/).

We just released LightlyTrain, a new open-source Python package (AGPL-3.0, free for research and educational purpose) for self-supervised pretraining of computer vision models: https://github.com/lightly-ai/lightly-train

Standard vision models pretrained on generic datasets like ImageNet or COCO often underperform on specific domains (e.g., medical, agriculture, autonomous driving). Fine-tuning helps, but performance is limited, and getting enough labeled data is expensive and slow.

LightlyTrain uses self-supervised learning (SSL) to pretrain models directly on your own unlabeled images or videos. This adapts the model to your specific visual domain before fine-tuning, leading to significantly better performance with less labeled data.

Key Features:

- No Labels Needed: Pretrain using your existing unlabeled image data.

- Better Performance: Consistently outperforms training from scratch and ImageNet-pretrained weights, especially in low-data regimes and domain-specific tasks (benchmarks in README/blog). We see gains across detection, classification, and segmentation.

- Domain Adaptation: Tailor models to your specific industry (manufacturing, healthcare, retail, etc.).

- Supports Popular Models: Works out-of-the-box with YOLO (v5-v12), RT-DETR, ResNet, ViTs, etc., integrating with frameworks like Ultralytics, TIMM, Torchvision.

- Easy to Use & Scalable: Simple pip install, minimal code to start, scales to millions of images, runs fully on-premise (single/multi-GPU). We built this because while SSL research is mature, making it easily accessible and effective for industry computer vision teams was hard. LightlyTrain aims to bridge that gap.

We’ve benchmarked it on COCO, BDD100K (driving), DeepLesion (medical), and DeepWeeds (agriculture), showing strong improvements over baselines (details in the repo/blog post linked below). For example, on COCO with only 10% labels, LightlyTrain pretraining boosted YOLOv8-s mAP by +14% over ImageNet weights and +34% over no pretraining.

- GitHub Repo: https://github.com/lightly-ai/lightly-train

- Docs: https://docs.lightly.ai/train

- Detailed Blog Post/Benchmarks: https://www.lightly.ai/blog/introducing-lightly-train

- Quick Demo Video: https://youtu.be/5Lmry1k_cA8

We’re here to answer any questions! Happy to discuss the tech, benchmarks, or use cases. Commercial licenses are also available for businesses needing different terms.

I’ve been playing around with SIMD since uni lectures about 10 years ago. Back then I started with OpenMP, then moved to x86 intrinsics with AVX. Lately I’ve been exploring portable SIMD for a side project where I’m (re)writing a Numpy-like library in Rust, mostly sticking to the standard library. Portable SIMD has been super helpful so far.

I’m on an M-series MacBook now but still want to target x86 as well, and without portable SIMD that would’ve been a headache.

If anyone’s curious, the project is here: https://github.com/IgorSusmelj/rustynum. It's just a learning exercise for learning Rust, but I’m having a lot of fun with it.

OpenAI O3-Mini 1 year ago

Does anyone know why GPT4 has knowledge cutoff December 2023 and all the other models (newer ones like 4o, O1, O3) seem to have knowledge cutoff October 2023? https://platform.openai.com/docs/models#o3-mini

I understand that keeping the same data and curating it might be beneficial. But it sounds odd to roll back in time with the knowledge cutoff. AFAIK, the only event that happened around that time was the start of the Gaza conflict.

I think the results show that just in general the compute is not used well. That the CPU took 8.4ms and GPU took 3.2ms shows a very small gap. I'd expect more like 10x - 20x difference here. I'd assume that the onnxruntime might be the issue. I think some hardware vendors just release the compute units without shipping proper support yet. Let's see how fast that will change.

Also, people often mistake the reason for an NPU is "speed". That's not correct. The whole point of the NPU is rather to focus on low power consumption. To focus on speed you'd need to get rid of the memory bottleneck. Then you end up designing your own ASIC with it's own memory. The NPUs we see in most devices are part of the SoC around the CPU to offload AI computations. It would be interesting to run this benchmark in a infinite loop for the three devices (CPU, NPU, GPU) and measure power consumption. I'd expect the NPU to be lowest and also best in terms of "ops/watt"

Is there somewhere an overview of the progress we made on the software side for training and inference of LLMs? It feels like we squeezed 10-100x more out of the hardware since llama appeared. This crazy progress will probably saturate though as we reach theoretical limits, no?

This is really cool.

Song: https://app.suno.ai/song/83680b6f-db37-44de-adf9-3f7fff6b79d...

Prompt: A 90s hip-hop song with a male singer with a deep voice singing about how AI models are creating new songs after being trained on all the data of artists. Talks about AI models are stealing the show.

Lyrics:

[Verse] Step back, my friend, 'cause the future's here AI models spittin' rhymes, oh so clear (oh so clear) Trained on data from all the greats Now they're droppin' beats that dominate (dominate)

[Verse 2] Don't need no ghostwriters or melody makers These AI models are the true risk-takers (oh yeah) Analyzin' every flow, every precise word Stealing the show, that's just absurd (it's absurd)

[Chorus] AI takin' over, breakin' the mold Stolen styles, but they're icy cold (they're icy cold) The game's been changed, no human control The rise of the AI flow, takin' its toll (takin' its toll)