HN user

mythz

10,569 karma

https://servicestack.net

Posts203
Comments1,633
View on HN
llmspy.org 6h ago

OSS ChatGPT WebUI v4 – Projects, Agent Profiles, Server Tools, 1-Click Sharing

mythz
1pts0
llmspy.org 1d ago

OSS ChatGPT WebUI v4 – Projects, Agent Profiles, Server Tools, 1-Click Sharing

mythz
2pts0
llmspy.org 2d ago

OSS ChatGPT WebUI v4 – Projects, Agent Profiles, Server Tools, Publishing

mythz
3pts0
about.fb.com 4mo ago

Meta and AMD Partner for Longterm AI Infrastructure Agreement

mythz
3pts0
llmspy.org 5mo ago

Llms.py new workspace editor for AI-powered browser automations

mythz
1pts0
llmspy.org 5mo ago

OSS ChatGPT WebUI – 530 Models, MCP, Tools, Gemini RAG, Image/Audio Gen

mythz
131pts32
llmspy.org 6mo ago

OSS ChatGPT WebUI – 530 Models, Tools, MCP, Gemini RAG, Image/Audio Gen

mythz
3pts0
llmspy.org 6mo ago

OSS ChatGPT WebUI – 530 Models, Tools, MCP, Gemini RAG, Image/Audio Gen

mythz
4pts0
llmspy.org 6mo ago

Show HN: llms.py OSS ChatGPT CLI and Web UI with Tool Calling, RAG, Extensions

mythz
2pts0
llmspy.org 6mo ago

llms .py – Extensible OSS ChatGPT UI, RAG, Tool Calling, Image/Audio Gen

mythz
2pts0
github.com 8mo ago

OSS Alternative to Open WebUI – ChatGPT-Like UI, API and CLI

mythz
101pts32
github.com 8mo ago

OSS Alternative to Open WebUI – ChatGPT-Like UI, API and CLI

mythz
9pts0
github.com 8mo ago

OSS ChatGPT-Like UI, API and CLI

mythz
1pts0
github.com 9mo ago

llms.py – Lightweight ChatGPT-Like API, UI and CLI

mythz
1pts0
github.com 9mo ago

llms.py – Lightweight ChatGPT-Like API, UI and CLI

mythz
4pts0
servicestack.net 9mo ago

Show HN: llms.py – Local OpenAI Chat UI, Client and Server

mythz
1pts0
servicestack.net 9mo ago

Show HN: llms.py – Local OpenAI Chat UI, Client and Server

mythz
3pts0
servicestack.net 9mo ago

Show HN: Llms.py – Local ChatGPT-Like UI and OpenAI Chat Server

mythz
2pts0
servicestack.net 9mo ago

Show HN: llms.py – Local ChatGPT-Like UI and OpenAI Chat Server

mythz
1pts0
servicestack.net 9mo ago

Llms.py – Local ChatGPT-Like UI and OpenAI Chat Server

mythz
1pts0
github.com 10mo ago

Llms.py – Lightweight Open AI Chat/Image/Audio Client and Server

mythz
1pts0
servicestack.net 1y ago

Open-Source AI Server – April 2025 Update

mythz
3pts0
servicestack.net 1y ago

AI Server – April 2025 Update

mythz
1pts0
servicestack.net 1y ago

Open-Source AI Server – April 2025 Update

mythz
1pts0
fedoramagazine.org 1y ago

Fedora Linux 42 Released

mythz
12pts0
servicestack.net 1y ago

Generate Blazor Admin CRUD Apps from a Text Prompt

mythz
1pts0
servicestack.net 1y ago

Generate Blazor Admin CRUD Apps with Text to Blazor

mythz
1pts0
servicestack.net 1y ago

Typed Open AI Chat and Ollama APIs in 11 Languages

mythz
1pts0
openai.servicestack.net 1y ago

Self Hosted AI Server Gateway for Ollama, ComfyUI and FFmpeg Servers

mythz
2pts0
litdb.dev 1y ago

Show HN: Typed, Expressive SQL-Like QueryBuilders/ORM for TypeScript/JS

mythz
3pts0

Antigravity is now split into 2 Apps:

'Antigravity' which is an agent-first editor layout where you don't see your code and just prompt it.

'Antigravity IDE' which uses the Windsurf/VS Code editor, which is still what I primarily use in my day-to-day.

I hope they never retire the IDE, I don't think I can get used to prompting an AI Agent without being able to see my code to help workout what needs to be done.

Always happy to see new Gemini releases as IMO Antigravity Pro 16.67/mo plan (Annual) is still the best plan available and have been pretty happy with Antigravity IDE.

If it wasn't for Gemini/Antigravity I'd have to go with a Max Claude plan, as it stands now I can get by with just a Claude Pro plan to get Opus when I need it, whilst using Antigravity as my day-to-day workhorse.

Unfortunately Gemini Flash became too expensive to use as a general purpose model (i.e. for AI features in Apps), luckily there are plenty of cheaper Chinese models to fill that gap now.

Buying at a high price just gives money to the early investors, the company itself doesn't get richer unless they're doing a raising capital round, but after an IPO you're just giving liquidity (+ profit) to early investors.

If you wait till the price comes down to its natural level, you'll be able to buy more of the company for less money.

Image Compression 1 month ago

Worth highlighting the QOI Image format (qoiformat.org) showing that you can also get significant image compression benefits with a simple 1-page specification [1]:

QOI is fast. It losslessly compresses images to a similar size of PNG, while offering 20x-50x faster encoding and 3x-4x faster decoding.

QOI is simple. The reference en-/decoder fits in about 300 lines of C. The file format specification is a single page PDF.

[1] https://qoiformat.org/qoi-specification.pdf

Because they want to optimize it for their models and don't want to be blocked by waiting for PRs to merge or be rejected.

There's plenty of reasons to start your own fork that you have full agency of, as long as the OSS License is maintained anyone will be able to benefit from any new features they want to make use of.

Grok 4.3 3 months ago

I said speed was great, Cerebas and Groq can provide better performance, likewise Fast versions of Cursor's Composer and Claude.

The reported speed like benchmarks is only a reported number on paper, we'll see how it holds up in real world usage, so far OpenRouter is only reporting 73tps

[1] https://openrouter.ai/x-ai/grok-4.3

Grok 4.3 3 months ago

Ok speed (202.7 tok/s) and value (1.25 -> 2.50) look great, with pretty decent intelligence.

Had an M4 iPad Pro for nearly 2 years. Everything's so fast and fluid I doubt I'd notice if the CPU were twice as fast.

If this ever died I'd likely replace it with an Air - the Pro is overkill for what's basically a consumption device.

It absolutely matters, especially when done in unison like this.

Cancelling ChatGPT sends a signal that you don't agree with weaponizing AI. Switching to Claude says you support Anthropic's principled stance against it. If you have a strong opinion either way, today is the day to vote with your wallet.

Dismissing every small action as meaningless is just apathy and how nothing ever changes.

These 2 Exceptions shouldn't have to be disputed.

At this point I'd go far to say I wouldn't trust any company with my AI history that caves to DoD demands for mass domestic surveillance or fully autonomous weapons.

Your AI will know more about you than any other company, not going to be trusting that to anyone who trades ethics for profits.

Servo isn't a JS engine. Do you mean why didn't they abandon their mission statement of developing a truly independent browser engine from scratch, abandon their C++ code base they spent the last 5 years building, accept a regression hit on WPT test coverage, so they can start hacking on a completely different complex foreign code-base they have no experience in, that another team is already developing?

I consider HuggingFace more "Open AI" than OpenAI - one of the few quiet heroes (along with Chinese OSS) helping bring on-premise AI to the masses.

I'm old enough to remember when traffic was expensive, so I've no idea how they've managed to offer free hosting for so many models. Hopefully it's backed by a sustainable business model, as the ecosystem would be meaningfully worse without them.

We still need good value hardware to run Kimi/GLM in-house, but at least we've got the weights and distribution sorted.

Sure he's still one of the top players, but he's not as strong this year and OP is suggesting he still has an edge against the GOAT, who this year:

- Has won Freestyle WC

- Has won SCC

- Has won 2x Titled Tuesday's

- Has won a Freestyle Friday

Hikaru can snipe a win off Magnus here and there, but I don't think there's any time control or format where he could win a long series of chess matches against Magnus.

Hikaru is either in a slump or his skill is starting to age: hasn't won Titled Tuesday since November, hasn't won Freestyle Friday this year, came last in Speed Chess Championship, etc.

We'll see how well he does in Candidates this year to see if he's still a top contender. Although I do believe this is his last chance to fight for the world title.

Really looked forward to this release as MiniMax M2.1 is currently my most used model thanks to it being fast, cheap and excellent at tool calling. Whilst I still use Antigravity + Claude for development, I reach for MiniMax first in my AI workflows, GLM for code tasks and Kimi K2.5 when deep English analysis is needed.

Not self-hosting yet, but I prefer using Chinese OSS models for AI workflows because of the potential to self-host in future if needed. Also using it to power my openclaw assistant since IMO it has the best balance of speed, quality and cost:

It costs just $1 to run the model continuously for an hour at 100 tokens/sec. At 50 tokens/sec, the cost drops to $0.30.

Doing inference with a Mac Mini to save money is more or less holding it wrong.

No one's running these large models on a Mac Mini.

Of course if you buy some overpriced Apple hardware it’s going to take years to break even.

Great, where can I find cheaper hardware that can run GLM 5's 745B or Kimi K2.5 1T models? Currently it requires 2x M3 Ultras (1TB VRAM) to run Kimi K2.5 at 24 tok/s [1] What are the better value alternatives?

[1] https://x.com/alexocheema/status/2016404573917683754

Did the napkin math on M3 Ultra ROI when DeepSeek V3 launched: at $0.70/2M tokens and 30 tps, a $10K M3 Ultra would take ~30 years of non-stop inference to break even - without even factoring in electricity. Clearly people aren't self-hosting to save money.

I've got a lite GLM sub $72/yr which would require 138 years to burn through the $10K M3 Ultra sticker price. Even GLM's highest cost Max tier (20x lite) at $720/yr would buy you ~14 years.

Not concerned with electricity cost - I have solar + battery with excess supply where most goes back to the grid for $0 compensation (AU special).

But I did the napkin math on M3 Ultra ROI when DeepSeek V3 launched: at $0.70/2M tokens and 30 tps, a $10K M3 Ultra would take ~30 years of non-stop inference to break even - without even factoring in electricity. You clearly don't self-host to save money. You do it to own your intelligence, keep your privacy, and not be reliant on a persistent internet connection.