HN user

Alifatisk

6,059 karma
Posts256
Comments3,144
View on HN
twitter.com 2d ago

The president of Argentina forgot to unschedule his victory tweet

Alifatisk
2pts3
arxiv.org 3d ago

Do Language Models Plan Ahead for Future Tokens? (2024)

Alifatisk
2pts0
www.the-independent.com 14d ago

"Most relaxing song" used to calm patients before surgery (2019)

Alifatisk
3pts0
www.techradar.com 25d ago

Cybersecurity company 360, "China's version of Mythos" unveils

Alifatisk
3pts0
slate.com 1mo ago

SmarterChild, Long before ChatGPT, a generation learned how to talk to machines

Alifatisk
3pts0
www.unibrands.co 1mo ago

Kuru Toga Dive: $100 mechanical pencil

Alifatisk
11pts12
detroitchinatown.org 1mo ago

Why the NATO Logo Is One of the Most Recognizable Symbols in the World

Alifatisk
2pts0
www.kimi.com 1mo ago

Moonshot AI releases Kimi WebBridge, a browser extension for AI agents

Alifatisk
2pts0
esengine.github.io 1mo ago

DeepSeek reasonix, DeepSeek native coding agent with high caching and low cost

Alifatisk
729pts288
ss64.com 1mo ago

Sips, macOS native scriptable image processing system

Alifatisk
2pts0
en.wikipedia.org 2mo ago

DjVu file format, alternative to PDF

Alifatisk
2pts0
www.oranchak.com 2mo ago

Can you crack the famous unsolved cipher, created by the Zodiac Killer?

Alifatisk
2pts0
www.citroenet.org.uk 2mo ago

Citroën metropolis concept car (2010)

Alifatisk
34pts20
www.minizinc.org 2mo ago

MiniZinc, constraint modelling language solve discrete optimisation problems

Alifatisk
64pts0
www.kimi.com 3mo ago

Kimi vendor verifier – verify accuracy of inference providers

Alifatisk
310pts33
openai.com 3mo ago

GPT‑2: 1.5B release (2019)

Alifatisk
1pts0
gh-v6.com 3mo ago

IPv6 GitHub Proxy

Alifatisk
2pts0
suchir.net 3mo ago

When does generative AI qualify for fair use? (2024) By previous OpenAI employee

Alifatisk
3pts0
www.youtube.com 3mo ago

Interview, Axios co-founder Mike Allen sits down with OpenAI CEO Sam Altman [video]

Alifatisk
3pts0
deadeclipse666.blogspot.com 3mo ago

Disclosing bluehammer exploit, vulnerability is still unpatched

Alifatisk
6pts2
twitter.com 3mo ago

Qwen-3.6-Plus is the first model to break 1T tokens processed in a day

Alifatisk
60pts23
vector-db-bench.kcores.com 3mo ago

GLM-5.1 tops Vector DB Benchmark

Alifatisk
4pts0
docs.z.ai 4mo ago

GLM-5-Turbo have been released, optimized for OpenClaw scenario

Alifatisk
2pts0
site.aignited.id 4mo ago

When do you get 2× Claude?

Alifatisk
1pts0
ronin-rb.dev 4mo ago

Ronin – A Security Toolkit

Alifatisk
2pts0
www.youtube.com 4mo ago

Minecraft mod "The Aether" announced new alpha release, Aether II [video]

Alifatisk
1pts0
github.com 4mo ago

Ruby -run, utilities to replace common Unix commands

Alifatisk
5pts0
www.youtube.com 4mo ago

This YouTube video is a drawing app

Alifatisk
1pts0
arxiv.org 5mo ago

Monolith – The research paper behind TikToks algorithm (2022)

Alifatisk
2pts0
www.youtube.com 5mo ago

Music Generation comes to Gemini [video]

Alifatisk
1pts0

Honestly, I don’t have any sympathy at all. Anthropic can complain all they want, but they seemed fine with pirating books. What Moonshot AI has done is to offer almost Fable 5 comparable performance at lower prices than Anthropic insane margins. This is what I call competition, which the Director seem to embrace.

This is what the Chinese always been good at. Take expensive innovation and streamline it to lower prices. But we are at a point where labs like Moonshot actually contributes a lot to the research field as well. They are pushing the innovation forward and squeezing the prices. Very well done.

Whats even weirder is the bizarre mechanisms Anthropic implemented to prevent distills which they had to sacrifice their customers for. They hid the internal CoT reasoning and returns summarizations instead. This made it difficult for users to trace things. They made Fable 5 silently switched over to Opus 4.8 if it detected blacklisted prompts (almost anything triggered this) to sabotage distills. And now, they are still complaining about distills? So their customers have gotten sacrificed over nothing.

Whats even weirder is the timeframe here, no way the Moonshot team managed to plan conduct a large scale distill, then pre-train, RL, fine-tune, benchmark, marketing and release to their platform since Fable 5 got whitelisted.

they developed a sophisticated internal platform to conduct large scale distillation

I am very curious about this and would love to learn more on how they did this. Wish we had more details. I know the team behind DeepSeek have also done clever things to distill too. I am aware of these ”transfer stations” that acts as a proxy, but I don’t think they are helpful in this case.

In other good news "the model has been trained to minimize refusals for beneficial uses.".

Otherwise, this news feels like a tiny incremental improvement on Gemini Flash series to make it more efficient with token usage, subagent and cost. Nothing big.

Regarding their benchmark scores on CyberGym, I wonder why they didn't compare their 3.5 Flash Cyber model with Fable 5. I mean they included Mythos and GPT-Cyber, so why not Fable 5 too?

They also mentioned Gemini 3.5 Pro is in testing and its about to become available very soon. Another thing maybe worth discussing is the announcement of pre-training Gemini 4. Sadly, not much technical details to discuss on. Many comments in here seem to mostly be about how Google is behind the others, but honestly, is it really worth the investment to be #1 in Artifical Analysis every week?

Over the past 48 hours, demand has pushed close to the limits of our current capacity. To protect the experience of existing subscribers, we're temporarily pausing new subscriptions and prioritizing compute for current members. Existing subscribed users are not affected.

Such a beautiful paragraph to read, a company that prioritizes their current customers and focus on keeping them satisfied instead of just focusing on fast growth.

If this is true, doesn't that mean the customers is getting served Fable 5 quality with the price of Deepseek V4 Pro? That's amazing! This team knows how to squeeze the prices.

Save GPT-5.5 3 days ago

This cannot be healthy? I am aware that these LLMs is getting really good. But why would anyone let themselves build an emotional attachment with a blackbox that can be taken away at any time? Just by having an emotional bond with a machine is already bad, but that's for another discussion maybe.

Qwen 3.8 3 days ago

I remember when they released Qwen 3.7 Plus and Max. These models behaved way different from all prior models, it became too verbose. It wrote multiple paragraphs just to answer my prompt instead of the usual concise and direct way responding to me. I didn't like that at all, and I know Gemini also had this behaviour with with the Flash series until I managed to reduce it a bit with personal instructions (in the settings on Gemini website).

I haven't tried Qwen 3.8 Max yet, looking forward to it. My hope is that its way less verbose. Another thing I experience with the Qwen models is that I do not trust their benchmark scores at all. Have anyone played with Qwen 3.8 Max and can share their experience? Which model it come close to? Sonnet 5? GLm-5? DS V4 Pro? Flash? Gemini 3.5 Flash?

Qwen 3.8 4 days ago

You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.

You can also try it out on Qwen chat, Its free.

A human can definitely be faster than these trillion parameter thinking models.

I’m thinking this can be true when the steps and muscle memory is known beforehand to the human. Otherwise we also have to stop, think and experiment for a while before we can proceed.

Deepsec 4 days ago

Is this perhaps related to CapitalOne VulnHunter?

In my case, I'll probably want to wait until other providers appear through OpenRouter and then I'll try to judge how much I trust them.

Keep in mind that the Moonshot team have identified multiple providers who configure their setup wrong which make their model perform worse than expected. This is why Moonshot created Kimi Verifier, but I guess its up to the provider if they want to do that.

Good that they are keeping it, Kimis way of speaking and conveying some sort of EQ is absolutely the best. The other models might be better at certain things, but nothing comes close to how good Kimi is at understanding language, emotions and reading the room in conversations.

I should maybe also mention that I have not used the later models like Opus or Fable, so my opinion might be a bit outdated.

When I remember that this site even showed Kimi having the highest score at one point https://eqbench.com

So you’re telling me, these people have workflows thats so tightly integrated to gemini-2.5-flash that no other model matches it’s performance? Really?

Have they really looked at all alternatives and found none to be a viable option?

I might have underestimated how good 2.5-flash was. I understand the issue with pricing though.

This is why I believe, for a company, to never be reliant on closed-weight models.

Java 27 already? I just learned about Java 26. But I’m not complaining, the JEPs that is getting introduced on every release are quite exciting features. I highly recommend following the Java official YouTube channel, they publish entertaining, yet informative videos/shorts about tips/tricks/features.

Grok 4.5 13 days ago

You really think such thing would occur? Why would someone spend their time on running these bots for the sake of defending from criticism?

What a fun article to read. This is such a cool showcase of the languages capabilities and standard library that comes with it. There is a gem named ”Ronin”, which is supposed to cover cases like this. But in your case, it doesn’t seem to be needed anyways.

That’s the point I think. Remember the controversy when Github Copilot came out? Not because where it got its data from, but because people didn’t feel like they wrote code anymore, they just tabbed the autocompleted snippets and was finished with a task in much shorter time.

No way, is Orion Browser available on other platforms as well now? Does it mean I can finally do tests for Safari (webkit) without owning an Apple product or paying for a vm? Incredible.

Claude Sonnet 5 22 days ago

Sure, but I think doing it this way allows them to later on say they were transparent about it. Completely hiding this would make it very difficult for them excuse when getting caught.