HN user

XCSme

2,907 karma

Building self-hosted web analytics.

Self-Hosted Analytics with Heatmaps and Session Recordings: https://www.uxwizz.com

WordPress Analytics: https://www.wplytic.com

X (Twitter): @XCSme

Posts69
Comments2,903
View on HN
aibenchy.com 1mo ago

Show HN: One hundred LLMs Generating a HTML/CSS Solar System

XCSme
5pts1
mariadb.org 1mo ago

MariaDB now has a DuckDB storage engine

XCSme
2pts0
aibenchy.com 1mo ago

SVG of a Hamster Playing Table-Tennis

XCSme
21pts18
news.ycombinator.com 2mo ago

Tell HN: Gemini 3.5 Flash breaks in stupid ways

XCSme
9pts4
www.uxwizz.com 2mo ago

Cisco Announces End of Life for Smartlook

XCSme
2pts1
news.ycombinator.com 4mo ago

Ask HN: Are MiniMax Models Scams?

XCSme
3pts2
aibenchy.com 4mo ago

Grok 4.20 brings minimal improvements over Grok-4.1-fast

XCSme
2pts1
twitter.com 4mo ago

Why Not Boost?

XCSme
1pts1
aibenchy.com 4mo ago

Show HN: AI Benchy – AI benchmarks and comparisons

XCSme
1pts0
www.uxwizz.com 4mo ago

PostHog now shares hashed emails of new users with Reddit and LinkedIn

XCSme
4pts2
aibenchy.com 5mo ago

Show HN: AIBenchy – Independent AI Leaderboard

XCSme
1pts1
softuts.com 6mo ago

Mailchimp Free Plan Changes

XCSme
2pts0
softuts.com 6mo ago

InvokeAI Commercial Platform Shuts Down, Open-Source Project Continues

XCSme
2pts2
www.uxwizz.com 6mo ago

Prevent others sending emails using your domain name

XCSme
1pts2
languagetool.org 7mo ago

LanguageTool browser extension is no longer free

XCSme
2pts1
news.ycombinator.com 7mo ago

Does anyone run ads successfully?

XCSme
3pts3
softuts.com 8mo ago

Hetzner Servers Benchmark

XCSme
4pts1
community.n8n.io 9mo ago

N8n added native persistent storage with DataTables

XCSme
174pts106
www.uxwizz.com 9mo ago

PostHog's $75M Series E: From Analytics to AI-Powered DevTools

XCSme
1pts0
twitter.com 9mo ago

Bolt v2

XCSme
2pts1
softuts.com 10mo ago

Docker Hub is down, no images can be pulled

XCSme
6pts1
softuts.com 12mo ago

One of the Weirdest Bugs

XCSme
1pts0
docs.uxwizz.com 1y ago

Import SQL data on an existing server

XCSme
2pts1
analyticsdir.com 1y ago

Compare all website analytics platforms

XCSme
1pts2
softuts.com 1y ago

Hetzner Servers Benchmarks

XCSme
3pts0
softuts.com 1y ago

Setup Postiz on Coolify

XCSme
2pts1
softuts.com 1y ago

Setting Up ChartBrew on Coolify

XCSme
3pts1
posthog.com 1y ago

PostHog raises $70M series D at almost $1B valuation

XCSme
3pts0
docs.uxwizz.com 1y ago

Self-Hosted Heatmaps/Recordings via Docker

XCSme
2pts1
n8n.io 1y ago

N8n – Flexible AI workflow automation for technical teams

XCSme
195pts99

The linked website shows this for me:

451: Unavailable due to legal reasons We recognize you are attempting to access this website from a country belonging to the European Economic Area (EEA) including the EU which enforces the General Data Protection Regulation (GDPR) and therefore access cannot be granted at this time. For any issues, contact hello@appenmedia.com or call 770-442-3278.

Thanks for the feedback, really good points!

There is some short info about the methodology here: https://aibenchy.com/methodology/

GPT-5.6 Sol on Low beats Fable Medium by 10% Fable is number 20

Fable loses a lot of points because it often refuses to answer questions. Asking a basic tool-usage challenge, Fable responded with refusal: "This request triggered restrictions on violative cyber content and was blocked under Anthropic's Usage Policy. To learn more, see https://platform.claude.com/docs/en/build-with-claude/refusa...." Even in practice, you ask Fable something trivial, and it refuses to respond. I think the score accurately represents how the model is behaving in real-world usage.

Gemini-3.6 Flash then beats them both Gemini models are the most intelligent overall. The tasks are not coding-only. Gemini excels in general knowledge and domain specific knowledge. Gemini models, even old ones, still top many charts on specific use-cases[0][1]. Depending on how you weigh those cases, the leaderboard order can vary quite drastically, as some models are very strong in some domains and weak in others.

You mention randomly selected questions. How does that work? Randomly selected, means I have manually created the questions/challenges to span across various domains and agentic surfaces. Questions vary from coding tasks, tool usage, trivia questions, chess puzzles, car-wash-like challenges and more.

With n=22 and binary pass/fail Each test is run 3 times, so in total we have 66 tasks. Also, apart from correct/wrong answer, the final score also includes the pass rate for each test (how many out of three attempts), how good the reasoning is (they have a hidden reasoning score where available) and other small factors. Also, some tests in some categories involve a series of tasks/requirements (i.e. implement this function, call it, do some processing on the result, combine the result with some built-in knowledge data, etc.).

I do agree that 22 tests isn't that much, and I'm slowly adding more, but even without the leaderboard part, the comparison feature is what's I think is most useful. You can see for the exact same tasks, which models do better, which do it faster, which cost less, etc.

_what_ you're testing and break that out by dimension There is a category breakdown for the test results, so you can see and which sort of tasks models fail.

Everything aside, when you manually ask a model to test its capabilities, I don't think it takes many questions to realise how good/bad that model is. Sometimes one prompt is enough, you ask it to do something, and see how it reasons about it, how fast it does it, how efficient the steps are and how good the result is. Yes, the performance may vary across tasks, but I'm pretty sure if you did a blind test with a chatbot, you could easily realise how good the model is in just a few questions/tasks.

I think no benchmark is perfect, mine is far from it, but it's simply another different, independent data-point. Apart from that, I made this for myself, and I'm using it myself. I don't trust that all popular benchmarks are not in the training data, and I think many benchmark the wrong things which don't correlate to how I use the models day-to-day myself. I just made the results publicly available, in case any one else benefits from it. I've probably spent thousands in LLM costs, and probably more than 100 hour building this, without benefiting in any way from it (outside of the joy of building it and me using it personally to compare models); as long as models cost stays reasonable, I'll keep building it and test new models as soon as they are released.

[0] https://x.com/browser_use/status/2079602472516264010/photo/1 [1]: https://artificialanalysis.ai/evaluations/mmmu-pro#mmmu-pro-...

They heavily boosted the amount of users in the past months (I myself got a business account because they had $1k free credits, and a personal one because of all the resets they give).

Now they are trying to quickly boost revenue with ads.

Next, IPO?

And two years before it's just silently plugging the product. And it will be considered "increased brand exposure", not ads...

Interesting, that was expected to happen at some point.

I didn't see any mention of prices, click rates, conversion rates compared to other type of ads, like search ads.

I assume it's going to be bidding based too?

I don't like to divulge tests, but one of them is a chess puzzle.

would be great if you could add GPT 5.6 Sol XHigh and Fable 5 High as well

I would like too, but I avoided them for several reasons:

1) Cost - this is a hobby project, those models would cost tens of dollars for each benchmark run, multiply this by tens or hundreds of models and ...

2) Time - the high models are already taking a really long answer to respond (5-10minutes per question). I run each question with 3 repeats (run the same test three times), so it would take 30 minutes per test. If I change my tests, methodology, or add a new test, it would take a really long time to run the benchmark. Also, I like having results immediately when a new model is released, now I can post within 30 minutes of a model's release the benchmark results.

3) High reasoning usually does WORSE on most tests - if you look at the leaderboard, it's sometimes counter-intuitive, but models with high or max reasoning usually do worse than medium and low. This is because the questions are quite targeted/direct, and the models overthink the question and miss the solution. Or the long thinking context makes them perform poorly. The generation tasks (SVGs/HTML animation) are usually better with longer reasoning, but short code fixes, trivia questions, puzzles, etc. are answered by low/med reasoning with more accuracy in general

Also, Fable is borderline un-testable, it refuses to answer many questions, so it scores poorly anyway.

Gemini scores 21/22 because it answers all tests, and it does them correctly, consistently. The only failed test is I think because it miscounted the lines in a file, when responding on which line the bug was in a code snippet.

Oh, and I've also added weights to different categories, so Coding and Tool usage categories influence the score more. This done both to better account for how most people are being used, and also to reduce Gemini's dominance in general/domain specific knowledge.

So yes, Gemini models are at the top, even if I actually (not proud of it) tried to make tests that actually favour other coding-focused models.

I have created various questions/tests and put the models through the same tests.

I record whether the answers are correct, and the generation stats (costs, latencies, tokens used, etc.).

I have no idea why the Gemini models do so well.

I have recently added new tests, whose sole purpose was to find some cases on which Gemini 3 Flash fails (I don't like cherry-picking models or tests, but I also find it strange Gemini Flash models leading in accuracy). I made a more complex coding/tool-usage test, that I expected it to fail, it did fail it once locally in my debug tests, but when I finalized the test and ran the entire testing suite for all models, somehow Gemini 3 Flash still got it right...

Gemini models are REALLY intelligent (and they are actually my favorite model to use via the chat app to ask questions), but they somehow fail in real-word coding tasks where they have to modify files, check results, debug, etc.

My tests harness provides a lot of mock data, and limits the number of actions a model can choose from. I am starting to think that maybe the models are not bad, just that the coding harness are not optimized for those type of models, and Google doesn't really provide their own "Codex".

In my tests, 3.6 Flash is NOT more token efficient, so it actually ends up costing more than 3.5 Flash, even with the output price reduction.

EDIT: It less less verbose in final output though, but it reasons more.

I assume the optimization comes when you have long-running tasks with many tool calls, and by reasoning more, it reduces the number of tool calls needed.

tl;dr: 3.6 flash is a bit smarter than 3.5 flash, but also a bit more expensive.

My results [0] put Gemini 3.6 Flash at the top.

3.6 Flash high has same $1.5 input price as 3.5 Flash, but output is cheaper from $9.0 to $7.5.

Google said 3.6 Flash is more token efficient, but in my tests it's actually LESS token efficient[1] than 3.5 Flash, so despite the output price reduction, it still costs more.

[0] https://aibenchy.com/compare/google-gemini-3-6-flash-medium/...

[1] https://aibenchy.com/compare/google-gemini-3-6-flash-high/go...

Qwen 3.8 3 days ago

I was not referring to the input/output price, but the cost of doing a specific tasks, in practice it is ~10x cheaper than GLM-5.2 for example, to accomplish the same task (for the tasks it can do).

I have been happily using DeepSeek V4 Flash for the last couple of months now. I tried GLM-5.2 for a while, but it was too slow and verbose compare to DeepSeek V4 Flash. If I have a basic skill I need to execute, DeepSeek V4 flash is still the best model for it.

Qwen 3.8 3 days ago

I use it through OpenRouter via Kilo Code VS Code extension.

You can check the logs in OpenRouter and see which providers it used and how many tokens you used.

Qwen 3.8 4 days ago

I am not even sure if Fable is as smart as they say, I can't get it to answer almost any question, it always refuses for "cyber-security" concerns...

Qwen 3.8 4 days ago

DeepSeek V4 pricing is insane, 10x-30x cheaper to use than most other models, and it usually is good enough for most tasks.

Benchmarks look ok, but they don't mention anything about the issue with the model being extremely slow and verbose.

That being said, it's awesome to have such an open-source model, even if now it's unusable mostly locally, with hardware improvements, in a couple of years, the verbosity/speed wouldn't matter as much as the intelligence.

I finished benchmarking[0] it, but it was not fun, it only supports (max) reasoning and the model is quite slow. Apart from a few requests timing out, it also has some issues with tool calling/response format schemas (Moonshot rejected tools.function.parameters with anyOf schema).

It also, for some reason failed to generate either of the 2 coding demos (hamster svg and solar system css animation).

Intelligence-wise, it's between GPT-5.6 Terra and GPT-5.6 Sol. It's ~30% better than Kimi K2.6, but a lot slower and more expensive.

[0] https://aibenchy.com/compare/moonshotai-kimi-k3-max/moonshot...

I am trying to benchmark it, but it only supports (max) reasoning, and even for simple questions, it takes forever to answer/times out :(

Only supporting "max" reasoning is weird, their parameters are quite inflexible atm:

    Important limits:

    reasoning_effort currently supports only max; K3 always has thinking mode enabled.

    max_completion_tokens defaults to 131072 and can be set up to 1048576.

    temperature=1.0, top_p=0.95, n=1, presence_penalty=0, and frequency_penalty=0 are fixed; omit them from requests.

    Return the complete assistant message unchanged in multi-turn conversations and tool calls.

    Vision input does not support public image URLs. Use base64 or ms://<file-id>, and make content an array of objects.

    Web search is being updated and is not recommended for production workflows in the near term.

Manually select a rectangle on a video frame, then do basic computer vision to detect notes, or even a simple image processing algorithm to find the lines and notes.

I was excited to use it, as I really wanted something like this, then I realized it needs AI/Claude (?)

That sounds like it can get quite costly. Probably there are ways to do it without AI, I would rather manually annotate the tab area with a visual editor.

With subscriptions, you want to have ways to increase the subscription amount and retain people, which usually leads to adding features no one asks for and bloating the product, trying to upsell users.