Yeah, this has been my progression as well. Maybe next it'll be just using plain pi when you figured out exactly what you want from omp and what you don't
HN user
schmorptron
Are thinking models only the reasonable tradeoff vs using much larger non thinking ones because the cost of output tokens is below that of input tokens?
That's a more than 2x jump in parameter count. I know it's not a measure of quality by itself, but it will be interesting how it "scales". Bust it looks like they're gonna be competing with the big boys now, pricing also approaches Gpt 5.6 Terra
yeah, but that's due to enterprise commitments that MS won't train on the user interactions
I don't think that's it for console manufacturers. They make the majority of their money on game sales, so they want the console itself to be used for as long as possible.
We're moving towards total surveillance slowly but surely. Age verification. Chat control. To an extent also the digital euro. It all seems hopeless, they're pushing this through despite what semblance of a democratic process we have clearly being against it. [that is not to speak of how undemocratic the european system is and how badly it needs reform. Von der leyen should never have been able to get the role she holds]
i love that word, and now it's genuinely (hehe) ruined. thanks, claude
I'm giving them the benefit of the doubt and interpreting it in a charitable way because they sound earnest about it, this is incredibly ambitious and cool-sounding, and I wish them all the best. It's something that's some sort of pipe dream, a noninvasive diagnosis machine that is able to use certain generic measurements and then derive insane levels of data from it. We've of course seen Theranos, but the holy grail remains.
Of course, there's always the tradeoff between research data collection and access vs user privacy, and striking that balance is incredibly hard. To make anything like this even remotely feasible you'll need a shitton of data and have it fully available to your researchers as well, while somehow safeguarding individual users. anonymizing medical data is impossible without rendering it near useless. Hoping they can figure that out! (Also, with human bodies being so different from one another, combatting bias is probably an eternal challenge)
you can build the datacenter right next to the tank and use the now-warm cooling water to pump into the tanks!
Cursor's composer models are finetuned kimi
It's "just" an opencode fork but it adds some nice features to try out while not being a full orchestrator metapackage like oh-my-opencode. Quite nice! Though it would be even nicer if this stuff came upstream or as an easy extension instead in the future
got one answer by reading the rest of the comments, makes sense that the diffusion process is inherently reasoning-like: https://www.inceptionlabs.ai/blog/introducing-mercury-2
What would a diffusing reasoning model look like? have a pre-defined length [thinking] block that gets diffused over a long time, and then the final output block uses what is in that thinking block as part of its input? And how do diffusion models decide the output length in the first place, is it a pre-set parameter? or does it diffuse an [end] token into the middle somewhere?
The irony of "we train on all of humanity's collective output, but god forbid anyone trains on ours" is still incredible
Cool project! I'll be trying it out. I've been a big fan of throwing whatever sources I have on a new topic i'm trying to get into into a llm "project" and then asking it to teach me, grounded on the actual content to speed things up.
But at the same time, I'm afraid getting everything laid out for you in exactly the way you want will erode some of the understanding you build by going through a primary source directly and figuring things out the hard way. So this having more focus on actually doing stuff by yourself seems right up my alley (while still tending to the LLM induced intellecutal laziness... ) .
Maybe this will replace raptor-mini as the "free" model on copilot plans? (but I don't see it at all yet on the student plan, in vscode or the cli)
the new intel ultra whatevername 3 series seems to come a bit closer there, so the framework pro with its explicit linux support might be an option
It's interesting that (for example for the explore agent https://github.com/Piebald-AI/claude-code-system-prompts/blo... ) they use a personality "you are a file search specialist" and "your strengths" framing. I thought that was largely thought to be useless, or even counterproductive nowadays? Does anyone know more about this stuff?
i see the reasoning traces in opencode (cli). maybe it's a setting?
I think part of it is also that we're able to still LARP as full developers of complex systems while vibe coding by seeing an interface that makes us look like l33t h4xx0rs even though we're just pressing continue 15 times
Associated paper: https://arxiv.org/html/2604.24827v1
for the current moment, Intel seems to be ahead of AMD for both power efficiency and iGPU performance. Panther Lake is really, really good.
In the gap between cost going down and profitability, is there not an increased risk of sybel attacks?
Oh, I was thinking more of user enters question into SO -> LLM answer on SO -> user evaluates whether LLM answer was sufficient (or system itself judges whether answer is also interesting to other users?) -> question + answer combo made public, judged by other users.
There are of course several huge issues with this, but thats why I prefaced it with ideal world hahaha
the biggest of which is why most users would want their questios publicized if the ChatGPT answer not on the stackoverflow platform will be enough or even better
Or how existing users and question-answering volunteers feel about just being cleanup and training data after LLMs
That's a hard one. SO's hostile community to newbies, like any expert community, comes from the longstanding users having seen the basic questions 1000s of times and understandably not wanting to answer variations of them over and over, while for the newbies those questions genuinely are there and they don't have the routine knowledge yet of where to look or how to even look for solutions in the first place.
In an ideal world, LLMs would take all of the basic RTFM style questions, and leave SO for the harder, but still general enough to be applicable to others-questions. LLMs seem to be getting pretty good at those as well though, so I don't know where that leaves us.
SO for discussions of taste? I have these two options to build this, how should i approach this? They tried to sell their own GPT wrapper for a while, didn't they? The use case I can see for that is: User asks question - LLM answers it - user is unsure about the answer - it gets posted as a SO thread and the rest of the userbase can nitpick or correct the LLM response.
Edit: I also seem to remember they had a job portal in the sidebar for a while, what happened to that? Seems like a reasonable revenue stream that is also useful to users.
I used a system prompt similar to this, where I just dumped the entirety of https://grugbrain.dev/ into it and prefaced it with the assistant having to emulate grug.
Didn't find it particularly useful, but is is funny!
Before this LLM age the solution would've been to make the user solve a leetcode problem to access a developer mode.
I actually feel like these integrations are fine, as long as they are opt-in or easily opt-outable of permanently. For now, I don't see the harm in adding another default search engine, it's much less obstrusive than the home page sponsored links. And if it gets them a little more independent from google by siphoning perplexity's seemingly infinite vc investment money, so be it.
I wonder if the rigidity could be improved while staying modular, maybe just use many more screws? I don't mind undoing more than 5 screws for the bottom to come off, make it 20 and it's still totally fine.