HN user

woadwarrior01

2,382 karma

I'm a SWE turned bootstrapped startup founder based in Dublin, Ireland. Was formerly working at FB, Reddit, Google and a few startups.

https://x.com/jeethu

Building: slopornot.app, privatellm.app and cleanlinks.app

Posts18
Comments845
View on HN
apps.apple.com 2mo ago

Show HN: I ported OmniAID image detection model to Apple's Neural Engine

woadwarrior01
4pts0
lite3.io 5mo ago

Lite³: A JSON-Compatible Zero-Copy Serialization Format

woadwarrior01
2pts0
apps.apple.com 11mo ago

Show HN: Clean Links – Free iOS app that removes tracking from URLs and QR codes

woadwarrior01
2pts0
arxiv.org 1y ago

GPT or BERT: why not both?

woadwarrior01
2pts0
news.ycombinator.com 2y ago

Show HN: The fastest way to run Mixtral 8x7B on Apple Silicon Macs

woadwarrior01
18pts22
huggingface.co 2y ago

Jamba-v0.1: An Apache 2.0 licensed 52B Mamba Transformer hybrid LLM base model

woadwarrior01
12pts0
news.ycombinator.com 3y ago

Show HN: I built an on-device LLM based chatbot for iPhones

woadwarrior01
20pts6
www.forbes.com 6y ago

Intel Lays Out Strategy for AI: It’s Habana

woadwarrior01
2pts0
repo.or.cz 7y ago

The mob account

woadwarrior01
1pts0
spa.mnesty.com 9y ago

Spamnesty – Have fun with spam

woadwarrior01
3pts0
rustyrazorblade.com 10y ago

RAMP Made Easy – Atomic reads in distributed databases

woadwarrior01
27pts0
nbaksalyar.github.io 10y ago

Rust in Detail: Writing Scalable Chat Service from Scratch

woadwarrior01
4pts0
nuitka.net 11y ago

Nuitka: a Python compiler

woadwarrior01
267pts135
www.youtube.com 11y ago

Docker Clustering on Mesos with Marathon

woadwarrior01
3pts0
ianmiell.github.io 11y ago

ShutIt – Complex Docker Deployments Made Simple

woadwarrior01
2pts0
darrennix.com 12y ago

How to choose the optimal domain name using ADWords split testing

woadwarrior01
1pts0
groups.google.com 14y ago

Algorithm for automatic cache invalidation

woadwarrior01
3pts1
jeethurao.com 17y ago

Tries and Ternary Search Trees in Python and Javascript

woadwarrior01
5pts0

Also, people seem to be oblivious to the existence of the neo-cloud: Nebius whose services and servers, a lot of western AI companies use. Nebius was spun out of Yandex 2 years ago to assuage the "supporting Russia" allegations and is still run and managed by an ex-Yandex crew. By that measure, I suppose using a lot of western AI tools is also equivalent to "supporting Russia".

Thanks for bringing this up. You're 100% right. But most people, even technical ones are oblivious to how much of a difference modern samplers and higher quality quantization algorithms make for on-device LLM inference and are stuck with good old top-p, top-k samplers and RTN quantization.

TBF mlx-vlm does support min-p sampling, but none of the other modern samplers that you list. Ollama and LM Studio are even worse with only top-p and top-k samplers.

Ironic that the app is named Nativ(e) and yet bundles a full Python runtime. Nonetheless, still less bloated than LM Studio, which bundles a full Python runtime and electron.js (which in turn bundles a whole browser runtime).

Works with the Bonsai-27B-1bit-mlx model on my mac.

Edit: this is running with MLX, and not with WebGPU on the browser, also it is a fail.

"""

You should *walk*.

Here’s why:

- *Distance*: 50 meters is very short — about 0.3 miles or 0.5 kilometers.

- *Time*: Walking takes roughly 2–3 minutes. Driving would take longer due to parking, starting the engine, and maneuvering.

- *Effort*: Walking is light exercise and avoids parking hassle.

- *Safety*: Less risk of misjudging parking space or damaging the car.

Unless you have a disability that makes walking difficult, walking is the most practical and efficient choice.

"""

Codex Resets 3 days ago

It's unlikely that a clear single winner will emerge in this competitive market.

In any case I am not sure pivoting from running local models to "cloud offering" (as in providing llm inference at their severs) is a sensible choice granted there is already competition in that space and they have no leverage there.

I agree. Incidentally, this is exactly what ollama are doing too.

Agree. OTOH, people have been using plotters for simulating handwriting. Also, nothing preventing someone from hand-copying an AI written piece as I'm sure lots kids are doing these days for take-home assignments.

Perf should be fairly straightforward to ballpark. You'll need to transfer roughly 2 . hidden_size . num_shards bytes over the network per token during autoregressive decoding. And divide that number by chunk size during prefill.

because Android re-skinned to use BUTTONS.

No. Steve's rage was justified, IMO. It was because Eric Schmidt was on Apple's board while simultaneously being Google's CEO and Google was surreptitiously building Android at the time. Mother of all conflict of interests.

There was a recent story that reminded me of it. Mike Krieger was on Figma's board and Anthropic's CPO, while Anthropic was surreptitiously building Claude Design.