I'd be curious about an.option that would allow glm use with a low end GPU like a 2080 ti...
HN user
aliljet
contact me here: pav.gup@gmail.com
The benchmarks here are confusing at best. Am I reading correctly that this model is essentially as good or better than all frontier models right now?
I was just using infinity parser 2 (flash, to be fair) for pennies self-hosted to run through thousands of pages of documents with remarkable confidence. I decided to use https://huggingface.co/datasets/allenai/olmOCR-bench to determine what was the best OCR tool, yesterday, but I've got no idea what the best is now. What is the dominant OCR eval right now? Between Baidu and Mistral this morning, I wonder if there's a new tool to switch to..
I'm curious about this. What models/tools have you been using?
How does this compare with infinty parser 2 which seemed to be running the table on every other OCR tool (https://huggingface.co/datasets/allenai/olmOCR-bench). To be fair, there's no single winning OCR benchmark and this isn't showing up anywhere yet..
This sounds incredible. Have these models effectively solved the problem of trying to use a fast-processing network to predict the world's state ahead? For example, to catch a ball?
The problem here is always the cost-benefit. For $200/mo, you're receiving subsidized best of breed access. There's no model competing for that price anywhere. If a 27B param model is what you choose, show me your hardware! I would love to be wrong...
Is this just one giant marketing plot?
Where can a user reasonably host this in an affordable way to access the local LLM revolution?
I'm really running into this deep at the edges of content creation. Take, for example, a need to general some kind of legal work. The cost of painstakingly checking and rechecking each case cited is reducing the value of these frontier models immensely.
Coding, however, is solved like magic. Easier to add tests, to be fair.
Is there a good benchmark tracking hallucinations? The models are all incredibly good now, even the open ones, and my hope is that the rate of hallucinations is something that's falling off in concert with larger and larger context lengths.
This is so cool. I would love to revitalize a generation of great, but perhaps boring older cars with FSD. Just so much work...
Why did Spirit die? Was there any last of this that had to do with their abysmal customer service?
What systems are you actively using? And what systems have you tried? It seems like law, generally, may be hitting a tipping point on LLM use...
This is a tough moment. Claude is simultaneously becoming substantially more expensive, substantially less reliable (single 9 of reliability), and substantially less performant. It's really hard to justify the cost of a subscription over there right now.
I wonder how this kind of response from Anthropic is actually being read by the community at large. If you consider the rough sentiment of the r/ClaudeCode subreddit against the r/Codex subreddit, you can see that there is a definite loudness among the folks departing ClaudeCode for Codex. Something big is shifting on the ground, I think.
Why is this being made public?
How can you reasonably try to get near frontier (even at all tps) on hardware you own? Maybe under 5k in cost?
Mythos is only real when it's actually available. If you're using Opus 4.7 right now, you know how incredibly nerfed the Opus autonomy is in service of perceived safety. I'm not so confident this will be as great as Anthropic wants us to believe..
I've found myself so deeply embedded in the Claude Max subscription that I'm worried about potentially makign a switch. How are people making sure they stay nimble enough not to get trarpped by one company's ecosystem over another? For what it's worth, Opus 4.7 has not been a step up and it's come with an enormously higher usage of the subscription Anthropic offers making the entire offering double worse.
How many levels of agents are here. Agents riding code by agents in a system driven by agents vibed by one lonely engineer in Redmond?
The real problem is that scientists doing this sort of early work more often than not want to burn hardware under their desks. Renting infrastructure in Google cloud isn't the only way...
I am hopeful that OpenAI will potentially offer clarity on their loss-leading subscription model. I'd prefer to know the real cost of a token from OpenAI as opposed to praying the venture-funded tokens will always be this cheap.
I've been looking for broken 3090s for a short while. And the whole market is funny. Most of these devices have had their VRAM and GPUs physically harvested (for no clear reason). The ones that are truly broken, however, are still trading in the ~$300 range making me think they're destined for more harvesting. Where are people buying these repairable GPUs? I'd gladly take the gamble for fun, honestly, less profit.
This is the rugpull that is starting to push me to reconsider my use of Claude subscriptions. The "free ride" part of this being funded as a loss leader is coming to a close. While we break away from Claude, my hope is that I can continue to send simple problems to very smart local llms (qwen 3.6, I see you) and reserve Claude for purely extreme problems appropriate for it's extreme price.
First, making sure to offer an upvote here. I happen to be VERY enthusiastic about local models, but I've found them to be incredibly hard to host, incredibly hard to harness, and, despite everything, remarkably powerful if you are willing to suffer really poor token/second performance...
What open models are truly competing with both Claude Code and Opus 4.7 (xhigh) at this stage?
This is the reality I'm seeing too. Does this mean that the subscriptions (5x, 10x, 20x) are essentially reduced in token-count by 20-30%?
I'm really curious about what competes with Claude Code to drive a local LLM like Qwen 3.6?
wait. that's insanity. where did you get those numbers from? the 5x plan is obviously the right place to be...