HN user

fzysingularity

194 karma

computer-vision, ml and all things meta.

Posts29
Comments113
View on HN
github.com 1mo ago

Omnigent: Meta-Harness for Coding Agents (Claude Code, Codex, Cursor, Pi)

fzysingularity
2pts0
huggingface.co 5mo ago

DeepSeek OCR 2: Visual Causal Flow

fzysingularity
2pts0
github.com 7mo ago

Unified Vision-Language Agents – Detect, Segment, OCR, Generate and More

fzysingularity
5pts1
chat.vlm.run 8mo ago

VLM Showdown: GPT vs. Gemini vs. Claude vs. Orion

fzysingularity
15pts1
chat.vlm.run 8mo ago

Show HN: Chat with Orion – a visual agent that sees, reasons and acts

fzysingularity
22pts10
twitter.com 8mo ago

ChatGPT uses YOLOv8 to detect UI elements

fzysingularity
1pts0
colab.research.google.com 1y ago

Build visual AI workflows from a prompt – OCR, detection, editing and more

fzysingularity
5pts4
docs.vlm.run 1y ago

How we solved multi-modal tool-calling in MCP agents – VLM Run MCP

fzysingularity
14pts6
docs.vlm.run 1y ago

video2json – Transcribe and analyze *hours-long* videos

fzysingularity
3pts1
twitter.com 1y ago

Georgi Gerganov on X: "x2 speed for WASM by optimizing SIMD" / X

fzysingularity
2pts0
twitter.com 1y ago

I'm Hooked on Devin ( Cognition_labs)

fzysingularity
1pts0
news.ycombinator.com 2y ago

Ask HN: GPT4V and function calling use-cases

fzysingularity
1pts0
twitter.com 2y ago

Silent voice messages transcribe to "Thanks for watching" on ChatGPT

fzysingularity
1pts0
news.ycombinator.com 2y ago

Ask HN: What are people using to automatically catalog images/video today?

fzysingularity
1pts0
news.ycombinator.com 2y ago

Ask HN: Which cloud provider offers AMD MI250/MI300?

fzysingularity
2pts5
docs.nos.run 2y ago

Serving LLMs on a Budget

fzysingularity
2pts0
twitter.com 2y ago

2024 predictions on the open-source LLM wars

fzysingularity
2pts2
news.ycombinator.com 2y ago

Fly.io for Multi-Cloud?

fzysingularity
1pts1
spillai.substack.com 2y ago

The AI Stack for the 3rd Epoch of Computing

fzysingularity
4pts0
erichartford.com 2y ago

My Own AI Server Cluster

fzysingularity
3pts0
github.com 2y ago

AGI-pack: Dockerfile generator for ML developers

fzysingularity
1pts1
intrinsic.ai 3y ago

Blog – Introducing Intrinsic Flowstate – Intrinsic

fzysingularity
1pts0
huyenchip.com 3y ago

Building LLM Applications for Production

fzysingularity
2pts0
www.trychroma.com 3y ago

Chroma

fzysingularity
1pts0
chat.openai.com 3y ago

OpenAI ChatGPT: Service Unavailable

fzysingularity
3pts0
twitter.com 4y ago

Benchmarking AWS Inferentia (inf1) chip performance

fzysingularity
2pts0
github.com 11y ago

OpenCV NumPy Converter Using Boost::Python

fzysingularity
2pts0
people.csail.mit.edu 11y ago

Learning Articulated Motions from Visual Demonstration

fzysingularity
3pts0
css.csail.mit.edu 12y ago

Bitcoin Transaction Graph Analysis

fzysingularity
1pts0

I think we all ought to look at the ZDR fine-print here.

I get that in principle that there's no retention, but these are powerful models that can comprehend, paraphrase and summarize your logs for the sake of "product" improvement. Who knows what's collected here.

It’s unclear to me what their desired outcome for a blog post like this. If you’ve ever worked in a robotics setting, 80% implies that 20% of your autonomous actions are incorrect. Imagine if this were the case for autonomous driving where your car misbehaves 1 in every 5 actions it takes.

Posts like this just reminds me of the end to end demos AV companies built in the early days using a single camera - only to realize that it’s harder than it looks years later into development.

That’s a pretty large binary for simply loading images.

In all honesty, opencv has stood the test of time and I’m certain newer LLMs will likely not attempt to rewrite it from scratch.

P.S. I’ve been a user since the IplImage days, circa 2007, and I’d still consider using it over most CV libraries today.

Claude Fable 5 1 month ago

I can’t help but think that there are so many astroturfed comments in here.

Seems like a concerted and distributed effort from the entire Anthropic team every time to get this on top of HN.

VLM Run (https://vlm.run) | 1x Product + 1x ML Staff Engineer | Santa Clara, CA (HQ)

We're building the inference and orchestration layer for production Vision-Language Models. We care deeply about fast and ergonomic visual inference, reliable structured outputs, and the observability to iterate on them.

A few things we've shipped recently you can poke at:

  1. Orion: our visual agent that reasons and acts over images, video, and documents. Chat at https://chat.vlm.run.
  2. mm-ctx: a Unix-style multimodal CLI (find, cat, grep, wc) that gives coding agents real context over images, video, and PDFs. Rust core, Python devex. 
  3. vlmbench:  single-file CLI for benchmarking VLM inference (TTFT, TPOT, throughput) across vLLM, Ollama, and SGLang.
Apply: https://app.dover.com/jobs/vlm-run

Email hiring "at" vlm.run with your GitHub + a couple recent projects.

[1] https://chat.vlm.run

[2] https://pypi.org/project/mm-ctx | https://www.vlm.run/open-source/mm

[3] https://github.com/vlm-run/vlmbench | https://www.vlm.run/open-source/vlmbench

The recent claude code leak also revealed that they're poisoning their competitors via anti-distillation policies baked in claude code CLI (fake tool calls, adding noise etc).

VLM Run (https://vlm.run) | 1x Infrastructure Engineer + 2x AI/ML Engineer | Santa Clara, CA (HQ)

VLM Run is building infrastructure for production Vision-Language Model (VLM) systems — fast inference, tool-use + orchestration, reliable structured outputs, and the observability to iterate quickly. We’re a deeply technical team of veteran AI / computer-vision engineers (20+ years combined, MIT/CMU PhDs) who’ve shipped production ML infrastructure across autonomous driving and LLMs.

Open roles:

1. Infrastructure Engineer (Full-time, ONSITE): $150K–$220K + 1–3% equity https://app.dover.com/apply/VLM%20Run/8d4fa3b1-5b38-42e1-927...

2. AI/ML Engineer (Full-time, ONSITE): $150K–$220K + 0.5–3% equity https://app.dover.com/apply/VLM%20Run/1a490851-1ea1-4f12-a0f...

Email hiring "at" vlm.run with your GitHub + a couple recent projects.

P.S. We recently launched Orion, our visual agent that can reason and act over images, videos and documents. You can chat with Orion at https://chat.vlm.run and see capabilities at https://docs.vlm.run.

Apply: https://app.dover.com/jobs/vlm-run

Real-time or continuous learning is great on paper, but to get this to work without extremely expensive regression testing and catastrophic forgetting is a real challenge.

Credit to the team for taking this on, but I’d be skeptical of announcements like this without at least 3–6 months of proven production deployments. Definitely curious how this plays out.

What do you think actually happened here in the past week?

They used Kimi, failed to acknowledge it in the original Composer announcement. Kimi team probably reached out and asked WTF? Their only recourse was to publicly disclose their whitepaper with Kimi mentioned to win brownie points about being open about their training pipeline, while placating the Kimi team.

VLM Run (https://vlm.run) | 1x Infrastructure Engineer + 2x AI/ML Engineer | Santa Clara, CA (HQ)

VLM Run is building infrastructure for production Vision-Language Model (VLM) systems — fast inference, tool-use + orchestration, reliable structured outputs, and the observability to iterate quickly. We’re a deeply technical team of veteran AI / computer-vision engineers (20+ years combined, MIT/CMU PhDs) who’ve shipped production ML infrastructure across autonomous driving and LLMs.

Open roles:

1. Infrastructure Engineer (Full-time, ONSITE): $150K–$220K + 0.5–3% equity https://app.dover.com/apply/VLM%20Run/8d4fa3b1-5b38-42e1-927...

2. AI/ML Engineer (Full-time, ONSITE): $150K–$220K + 0.5–3% equity https://app.dover.com/apply/VLM%20Run/1a490851-1ea1-4f12-a0f...

Email hiring "at" vlm.run with your GitHub + a couple recent projects.

P.S. We recently launched Orion, our visual agent that can reason and act over images, videos and documents. You can chat with Orion at https://chat.vlm.run and see capabilities at https://docs.vlm.run.

Apply: https://app.dover.com/jobs/vlm-run

Hugging Face Skills 5 months ago

uvx probably is the way to go here (fully self-contained environment for each skill), and use stdout as the I/O bridge between skills.

ELO scores for OCR don't really make much sense - it's trying to reduce accuracy to a single voting score without any real quality-control on the reviewer/judge.

I think a more accurate reflection of the current state of comparisons would be a real-world benchmark with messy/complex docs across industries, languages.

VLM Run (https://vlm.run) | Infrastructure Engineer + DevRel + AI/ML Engineer | Santa Clara, CA (HQ)

VLM Run is building infrastructure for production Vision-Language Model (VLM) systems — fast inference, tool-use + orchestration, reliable structured outputs, and the observability to iterate quickly. We’re a deeply technical team of veteran AI / computer-vision engineers (20+ years combined) who’ve shipped production ML infrastructure across autonomous driving and LLMs.

Open roles:

1. Infrastructure Engineer (Full-time, ONSITE): $150K–$220K + 0.5–3% equity https://app.dover.com/apply/VLM%20Run/8d4fa3b1-5b38-42e1-927...

2. Founding DevRel (Full-time, ONSITE/REMOTE): $90K–$140K + 0.5–3% equity https://app.dover.com/apply/VLM%20Run/de84c63e-fd0a-418b-929...

3. AI/ML Engineer (Full-time, ONSITE): $150K–$220K + 0.5–3% equity https://app.dover.com/apply/VLM%20Run/1a490851-1ea1-4f12-a0f...

Email hiring "at" vlm.run with your GitHub + a couple recent projects.

P.S. We recently launched *Orion*, our visual agent that can reason and act over images, videos and documents. You can chat with Orion at https://chat.vlm.run and see capabilities at https://docs.vlm.run.

Apply: https://app.dover.com/jobs/vlm-run

It's like going to the grocery store and buying tabloids, pretending they're scientific journals.

This is pure gold. I've always found this approach of evals on a moving-target via consensus broken.

I'd love to see Claude Code remove more lines than it added TBH.

There's a ton of cruft in code that humans are less inclined to remove because it just works, but imagine having LLM doing the clean up work instead of the generation work.