HN user

markab21

177 karma
Posts0
Comments60
View on HN
No posts found.

I'm mildly surprised that more people aren't using Nemo models for this reason. We've moved most of our processing to a combination of Nemo Ultra and Super, with some support for multi-model-specific tasks on Omni. The setup is working REALLY well for us, and I'm comfortable with the more measured pace of improvements. We work with many long-context problems, and the ecosystem is great.

There were a number of use cases where we needed to use Gemini (audio modality), and Ultra has been a VERY cost-effective alternative once we got through the nuances.

I'll pipe in here as someone working on an agentic harness project using mastra as the harness.

Nemotron3-super is, without question, my favorite model now for my agentic use cases. The closest model I would compare it to, in vibe and feel, is the Qwen family but this thing has an ability to hold attention through complicated (often noisy) agentic environments and I'm sometimes finding myself checking that i'm not on a frontier model.

I now just rent a Dual B6000 on a full-time basis for myself for all my stuff; this is the backbone of my "base" agentic workload, and I only step up to stronger models in rare situations in my pipelines.

The biggest thing with this model, I've found, is just making sure my environment is set up correctly; the temps and templates need to be exactly right. I've had hit-or-miss with OpenRouter. But running this model on a B6000 from Vast with a native NVFP4 model weight from Nvidia, it's really good. (2500 peak tokens/sec on that setup) batching. about 100/s 1-request, 250k context. :)

I can run on a single B6000 up to about 120k context reliably but really this thing SCREAMS on a dual-b6000. (I'm close to just ordering a couple for myself it's working so well).

Good luck .. (Sometimes I feel like I'm the crazy guy in the woods loving this model so much, I'm not sure why more people aren't jumping on it..)

I think the entire premise that the prompting is the surface area for optimizing the application is fundamentally the wrong framing, in the same way that in 1998 better cpam will save CGI. It's solving the wrong problems now, and the limitations in context and model intelligence require a tool like Dspy.

The only thing I'd grab dspy for at this point is to automate the edges of the agentic pipeline that could be improved with RL patterns. But if that is true, you're really shorting yourself by giving your domain DSPY. You should be building your own RL learning loops.

My experience: If you find yourself reaching for a tool like Dspy, you might be sitting on a scenario where reinforcement learning approaches would help even further up the stack than your prompts, and you're probably missing where the real optimization win is. (Think bigger)

Gemini 3.1 Pro 5 months ago

You just articulated why I struggle to personally connect with Gemini. It feels so unrelatable and exhausting to read its output. I prefer to read Opus/Deepseek/GLM over Gemini, Qwen and the open source GPT models. Maybe it is RLHF that is creating my distaste from using it. (I pay for Gemini; I should be using it more... but the outputs just bug me and feel more work to get actionable insight.)

Qwen3-Coder-Next 6 months ago

It's getting a lot easier to do this using sub-agents with tools in Claude. I have a fleet of Mastra agents (TypeScript). I use those agents inside my project as CLI tools to do repetitive tasks that gobble tokens such as scanning code, web search, library search, and even SourceGraph traversal.

Overall, it's allowed me to maintain more consistent workflows as I'm less dependent on Opus. Now that Mastra has introduced the concept of Workspaces, which allow for more agentic development, this approach has become even more powerful.

I love where you're going with this. In my experience it's not about a different persona, it's about constantly considering context that triggers, different activations enhance a different outcome. You can achieve the same thing, of course by switching to an agent with a separate persona, but you can also get it simply by injecting new context, or forcing the agent to consider something new. I feel like this concept gets cargo-culted a little bit.

I personally have moved to a pattern where i use mastra-agents in my project to achieve this. I've slowly shifted the bulk of the code research and web research to my internal tools (built with small typescript agents).. I can now really easily bounce between different tools such as claude, codex, opencode and my coding tools are spending more time orchestrating work than doing the work themselves.

Apps SDK 10 months ago

The skepticism is understandable given the trajectory of GPTs and custom instructions, but there's a meaningful technical difference here: the Apps SDK is built on the Model Context Protocol (MCP), which is an open specification rather than a proprietary format.

MCP standardizes how LLM clients connect to external tools—defining wire formats, authentication flows, and metadata schemas. This means apps you build aren't inherently ChatGPT-specific; they're MCP servers that could work with any MCP-compatible client. The protocol is transport-agnostic and self-describing, with official Python and TypeScript SDKs already available.

That said, the "build our platform" criticism isn't entirely off base. While the protocol is open, practical adoption still depends heavily on ChatGPT's distribution and whether other LLM providers actually implement MCP clients. The real test will be whether this becomes a genuine cross-platform standard or just another way to contribute to OpenAI's ecosystem.

The technical primitives (tool discovery, structured content return, embedded UI resources) are solid and address real integration problems. Whether it succeeds likely depends more on ecosystem dynamics than technical merit.

I've found myself more and more using local models rather than ChatGPT; it was pretty trivial to set up Ollama+Ollama-WebUI, which is shockingly good.

I'm so tired of arguing with ChatGPT (or what was Bard) to even get simple things done. SOLAR-10B or Mistral works just fine for my use cases, and I've wired up a direct connection to Fireworks/OpenRouter/Together for the occasion I need anything more than what will run on my local hardware. (mixtral MOE, 70B code/chat models)

For Llama-based progress - Reddit - /r/LocalLlama has been my top source of info, although it's been getting a little more noisy lately.

I also hang out on a few Discord servers: - Nous Research - TogetherAI / Fireworks / Openrouter - LangChain - TheBloke AI - Mistral AI

These, along with a couple of newsletters, basically keep a pulse on things.

Linux distributions, including Debian, offer a variety of desktop environments, each with its own design philosophy and user experience. If one environment doesn't suit your preferences, others might be more to your liking. It's worth exploring different desktop environments to find one that aligns with your expectations.

Debian and Ubuntu have similarities, but keep in mind - Ubuntu is derived from Debian, not the other way around. However, they differ in areas like release cycles, package management, and default configurations.

I used it as a consultant on a development project to help me organize some of the milestones and design goals in some documentation.

It wasn't that I didn't know the stuff, I do, but more helpful with quickly organizing and presenting information in a clean and well-written way. I did have to go through and re-write parts of it specific to our domain.. but it saved me many hours of work doing tedious organization of data.

I also tested it with helping create some SOP's for a new position in our very small company, even breaking down the expected tasks into daily schedules.

It's not that it's perfect, but it generates a bit of a boiler-plate starting point for me which then I can work with from there.

The claim is dead-right.

I own a Tesla Model 3, my wife drives a BMW i3, my daughter has a leaf.

The ONLY car we can effectively travel outside of the greater Tampa area without major headache is the Tesla.

The ONLY car that I would try to drive to New York from Florida in is the Tesla. (Yes, we've done it.. but would only try it in the Tesla)

As someone who regularly pilots a piston-single, I'm looking forward to what hybrid or full electric can do for us little guys. The takeoff phase of flight is a high-risk situation in events such as fuel contamination which could prove fatal.

Having a small buffer, even a few minutes of battery power to rely on while trying to get back to the field to land the "impossible turn" [1] would make me feel a lot better and could be the difference between life and death.

There are various groups (Pipstrel[2], Diamond[3]) that I'm aware of that are working on electric GA aircraft. For young pilots that are looking to train, the cost of jumping in a Cessna 172, the gold standard in GA trainers will cost at best $120-200/hr. Electric costs should be 1/5 (or better) of that in reality due to the absolute bargain of replacing a TBO electric engine, scheduled maintenance and relatively low level of complexity.

[1] https://www.aopa.org/training-and-safety/air-safety-institut...

[2] https://www.pipistrel-usa.com/electric-propulsion/

[3] https://www.flyingmag.com/diamond-da40-hybrid-electric-proto...

I'd like to get one but it doesn't fit my use case (namely long road trips in the summer).

I own a P3D, facing the same issues of long road trips 1-2x a year I concluded that I'll just spend the few hundred dollars and rent a car for the deep edge cases of my driving and the other 99% of my time I'll enjoy driving my Tesla.

99% of my normal day to day driving I just charge at home at night, I've used the supercharger network on a long road trip and it was surprisingly little-hassle.

I've noticed my dog (black lab) doing similar behavior consistently in the mornings, also after she had her breakfast. She'll walk away from the kitchen and her dog bowl with her tail wagging and doing some kind of sneezing sniffing-sound, also when we come home from being out of the house she'll do it too.

It's cute as hell.

Interesting.

Do you have any thoughts on composite-based aircraft like a Cirrus and flutter? A few times I've had a Cirrus SR22 into a pretty steep descent with poor controller sequencing for an approach into busy terminal space and had to push it down, but the plane felt solid even at 180-190kts TAS. I backed it off only because I get nervous with any unexpected turbulence which is not uncommon in Florida.

The Piper Saratoga I flew for a bit didn't seem to like the speed as much, that or the toga was a bit more vocal than the Cirrus in what it was feeling with regards to airspeed.