HN user

dazzaji

158 karma

Click here: https://news.ycombinator.com/user?id=dazzaji

[ my public key: https://keybase.io/civics; my proof: https://keybase.io/civics/sigs/QCufNppvl0j1drs4yNOjS65VhZb8NrnKU0CAPwGSNUM ]

Posts6
Comments78
View on HN

Nope, not bored at all. I'm clicking on most HN AI threads out of sheer curiosity, eager to learn new techniques and see what folks are thinking or building. Given its transformative ripple effects, AI feels like the single biggest shift reshaping the economy and society right now. Kind of the opposite of boring.

I’m pleasantly surprised this was AI assisted so deeply that inconsistencies like that slipped by you. The writing is really extraordinary. It made me want to read for fun again for the first time in decades. Thank you!

This discussion hits close to home. A few of us at Stanford and Consumer Reports have been working on a project called Loyal Agents (loyalagents.org ) that’s focused on the same core issue raised in the Economist article, namely how to make sure AI agents actually act in the interest of the people they represent.

The idea is to define what “loyalty” means for an AI agent in both technical and legal terms, and then build systems that can prove they’re acting on a user’s behalf (ie not a platform’s or advertiser’s).

It’s early-stage research, but the overlap with many of the questions here is striking. Would be great to get feedback from this crowd as the work evolves.

I’m part of the group working on Loyal Agents and happy to discuss it.

Here’s what Claude Sonnet 4.5 suggested to take this piece from something that sounds impressive but lacks substance to something that could actually deliver on its promise. I did this thought exercise to explore whether being AI-generated necessarily precludes brilliance. You be the judge - I think Claude succeeded in mapping the gap between the current draft and what a truly excellent version would actually require.

https://claude.ai/share/46dd4b7e-9adf-473d-8372-22cb1ae34249

Not quite - as I understand it box-counting measures global space-filling, manifolds handle local coordinate structure. Consider that the Earth is locally flat but globally spherical, and a Möbius strip vs cylinder are locally identical but globally different. Related problems, but the tools reveal different aspects of geometry. So I think whether “this is exactly what topological manifolds are for” depends what you’re trying to understand.

Synthi is an open web tool that instantly summarizes and synthesizes Hacker News threads and their linked articles, grouping every point of view by topic.

[Live demo here: https://prototypejam.github.io/synthesize/](https://prototypejam.github.io/synthesize/)

Why I Built This:

I love the deep discussions on HN, but I never have time to read a long article and a 400-comment thread. The tipping point for building this was when I nearly burned through my monthly API credits on another service just trying to synthesize a few threads! I needed my own tool.

How It Works:

1. Paste any HN thread URL into Synthi. 2. It instantly detects it and fetches the linked article. 3. Click "Full Analysis" for a unified, topic-based synthesis (with attribution to the article or the commenter). 4. Export the result, bookmark it, or listen to it (the output is text-to-speech friendly).

Key Features:

* Smart HN workflow: Auto-detects HN links for a one-click analysis. * Works with any URLs: Can also synthesize any two articles or pieces of text. * 100% client-side: All processing happens in your browser. No backend, no tracking. * Open source (MIT License): The code is yours to inspect, fork, and use.

The Code: [https://github.com/prototypejam/synthesize](https://github.com/prototypejam/synthesize)

A Note on the API Key:

Synthi is a "Bring Your Own Key" app. You'll need a Google Gemini API key, which is stored securely in your browser's local storage and never sent anywhere else.

For now, it's Gemini only, largely because the free tier on Google AI Studio is incredibly generous and a great way to get started. I'm considering adding support for other models like Claude or OpenAI via OpenRouter in the future.

I've been using this every day and I hope it's useful to some of you too. All feedback is welcome!

Wow - I'd forgotten all about this but just realized I have posts from an entire phase of earlier professional life - topic by topic and event by event - on an old blog there. Amazingly the browser remembered my login so I was able to find the URL. It's been quite a trip down memory lane revisiting some of the posts. Not sure I need to keep any of that published but I'll at least scrape and store it somewhere for old times sake. Maybe I'll find some buried gem of an idea when I scan them during the great scrape. Or - optimistically - perhaps a future zillion-token context LLM will uncover some personal patterns that unleash deep and actionable insights. Irrespective of the measurable value, I just hate to see the old posts dissapear forever.

This is a fun project to be sure. I just wish the author would not refer to the experiment as an "autonomous startup builder" unless they mean it humorously. Having poked around the GitHub repo and read through the materials, it seems like more of an AI coding assistant running in a loop that built and deployed a broken web application with no users, no business model, and no understanding of what problem it was trying to solve. There were quasi-autonomous processes and there were things that were built, but nothing I'd call a startup.

One of the things that I found most frustrating about USB-C hubs is how hard it is to find one that actually gives you multiple USB-C ports. I have several USB-C devices but most hubs just give you one USB-C port and a bunch of USB-A ports. At most it’s 2 USB-C ports but only with the hub that plugs into both USB-C ports on my MacBook Pro (so I’m never able to get more ports than I started with). The result is I end up having to keep swapping devices. For a connector that was supposed to be the "one universal port," it's weird that most hubs assume you only need one USB-C connection. Has anyone found a decent hub with multiple USB-C data outputs?

That system would make a tidy startup, especially if tightly integrated with an open source office suite behind the scenes (LibreOffice, OpenOffice, etc) and a generative AI native UX.

Got it, thanks for clarifying! So if I’m understanding you right, you’re saying that all the generative stuff the LLM does—like creating text—basically becomes part of the ‘arguments’ the original post talks about, and then that gets paired with a tool call (like inserting into a text editor, doing edits, etc.). I was focused on the tool call not the argument content aspect of the post.

And it sounds like you’ve had a lot of success with this approach in an impressive variety of application types. May I ask what tooling you usually use for this (eg custom python for each hack? MCP? some agent framework like LangGraph/ADK/etc, other?)

I’m still stuck on the first sentence "An LLM should never output anything but tool calls and their arguments” because it just doesn’t make sense to me.

Tool calling is great, but LLMs are - and should be used as - more than just tool callers. I mean, some tools will have to be other LLMs doing what they’re good at, like writing a novel, summarizing, brainstorming ideas, or explaining complex topics. Tools are useful, but the stuff LLMs actually do is also useful. The basic premise that LLMs should never output anything beyond tools and arguments is leaving most of the value of LLMs on the table.

No one actually pays that price. The $1000 misrepresented in the article…

With respect, that is absolutely incorrect. People absolutely pay over $1000 and do so monthly. For example, Kaiser of Northern California makes it very difficult for their doctors to prescribe these, and nearly impossible to get a prescription for Monjaro (which is particularly effective). Therefore, Kaiser patients/insured for whom these drugs are of immense benefit but who must have their prescriptions from out of network physicians receive ZERO insurance coverage. This means they get neither the negotiated insurance price discount nor any co-pay on the full cost. I am directly aware of this. And it is a travesty. Yet the benefits of these drugs is so significant and uniquely available through these drugs that in a sense, if it is possible to pay, then pay one must. Because in effect they are invaluable.

I was fortunate to get early access to the new Agent SDK and APIs that OpenAI dropped today and made an open source project to show some of the capabilities [1]. If you are using any of the other agent frameworks like LangGraph/LangChain, AutoGen, Crew, etc I definitely suggest giving this agent SDK a spin.

To ease into it, I added the entire SDK with examples and full documentation as a single text file in my repo [2] so you can quickly get up to speed be adding it to a prompt and just asking about it or getting some quick start code to play around with.

The code in my repo is very modular so you can try implementing any module using one of the other frameworks to do a head-to-head.

Here’s a blog post with some more thoughts on this SDK [3] and some if its major capabilities.

I’m liking it. A lot!

[1] https://github.com/dazzaji/agento6

[2] https://raw.githubusercontent.com/dazzaji/agento6/refs/heads...

[3] https://www.dazzagreenwood.com/p/unleashing-creativity-with-...

Here’s the conclusion of a much more refined initial review by Andrej Karpathy [1] which, I think overall, comports with the substance of my own hot take:

“As far as a quick vibe check over ~2 hours this morning, Grok 3 + Thinking feels somewhere around the state of the art territory of OpenAI's strongest models (o1-pro, $200/month), and slightly better than DeepSeek-R1 and Gemini 2.0 Flash Thinking. Which is quite incredible considering that the team started from scratch ~1 year ago, this timescale to state of the art territory is unprecedented. Do also keep in mind the caveats - the models are stochastic and may give slightly different answers each time, and it is very early, so we'll have to wait for a lot more evaluations over a period of the next few days/weeks. The early LM arena results look quite encouraging indeed. For now, big congrats to the xAI team, they clearly have huge velocity and momentum and I am excited to add Grok 3 to my "LLM council" and hear what it thinks going forward.”

[1] Full review at: https://x.com/karpathy/status/1891720635363254772?s=46&t=91u...

That sounds like very evocative prose. Would you be up for sharing some of that fiction? I haven’t tried Grok 3 for that purpose and now I’m curious.

It’s gracious of you to say that you’d be sorry, and I did run my comment through 4o (perhaps ironically) which caught a slew of typos and weird grammar issues and offered some improvements. But the robotic sound and anything else you don’t like are my own responsibility. Do you, perhaps, have any thoughts on the substance of the comment?

My Time at MIT 1 year ago

I started my affiliation with MIT in 1997 as a lecturer on eCommerce Architecture—a topic that felt both exciting and exotic back then, as the internet was just beginning to transform the world. Walking through those doors for the first time, I was immediately struck by the brilliance of the people around me. Virtually everyone I met seemed like the smartest person I’d ever encountered, and I couldn’t help but feel like an imposter at times. But that feeling also came with the thrill of being part of a community that constantly challenged and elevated me.

In 2007, I moved over to the Media Lab to focus on computational law, a field that thrives on MIT’s interdisciplinary ethos. Over the years, I’ve had the privilege of collaborating with CSAIL researchers from time to time. These experiences were always a blast—not just because of the cutting-edge research, but also because navigating the Frank Gehry-designed Stata Center was its own adventure. I loved spelunking through its quirky passageways, stumbling across obscure treasures like tucked-away whiteboards filled with half-finished equations or discovering yet another coffee machine in a corner I hadn’t visited before.

Since the pandemic, I’ve been living back in the Bay Area, so much of my MIT involvement is now remote. Yet even from a distance, I remain in awe of the people, ideas, and unrelenting creativity that define the Institute. Reading this reflection brought back so many memories—and a wave of nostalgia. It’s inspiring to see how others have experienced their time at MIT, and it makes me want to book a visit back to the campus just to soak it all in again. There’s something about MIT that stays with you, no matter how far away you are.

This article seems to be mostly about the slowdown in globalization since 2008, driven by economic uncertainty and advancements in automation reshaping global value chains. It makes sense that reshoring is becoming more viable for industries with higher levels of automation and that control and risk reduction are taking precedence over cost savings. I suppose a really good model can show greater control and risk reduction represent more value than the rewards of direct cost savings.

I imagine this trend will only accelerate with the renewed focus on tariffs initiated in the United States, which seem to be rippling across the globe. Tariffs not only increase costs but also encourage nations and companies to rethink dependencies on international supply chains. I'm no economist, but my take away is that the tariff wars, combined with the rise of automation, might just be catalyzing a structural shift toward more localized production on a global scale. That wouldn’t necessarily be so bad.

I rely on ElevenReader several times a week for quick text to voice on snippets of text I’m working on or sometimes on full web pages when I hand it a url. It’s quick and easy to use and the performance and quality is high.

Late Sunday night, I gained access to OpenAI’s newly launched Deep Research and immediately tested it on a draft blog post about Uniform Electronic Transactions Act (UETA) compliance and AI-agent error handling [1]. Here’s what I found:

Within minutes, it generated a detailed, well-cited research report that significantly expanded my original analysis, covering: * Legal precedents & case law interpretations (including a nuanced breakdown of UETA Section 10). * Comparative international frameworks (EU, UK, Canada). * Real-world technical implementations (Stripe’s AI-driven transaction handling). * Industry perspectives & business impact (trust, risk allocation, compliance). * Emerging regulatory standards (EU AI Act, FTC oversight, ISO/NIST AI governance).

What stood out most was its ability to: - Synthesize complex legal, business, and technical concepts into clear, actionable insights. - Connect legal frameworks, industry trends, and real-world case studies. - Maintain a business-first focus, emphasizing practical benefits. - Integrate 2024 developments with historical context for a deeper analysis.

The depth and coherence of the output were comparable to what I would expect from a team of domain experts—but delivered in a fraction of the time.

From the announcement: Deep Research leverages OpenAI’s next-generation model, optimized for multi-step research, reasoning, and synthesis. It has already set new performance benchmarks, achieving 26.6% accuracy on Humanity’s Last Exam (the highest of any OpenAI model) and a 72.57% average accuracy on the GAIA Benchmark, demonstrating advanced reasoning and research capabilities.

Currently available to Pro users (with up to 100 queries per month), it will soon expand to Plus and Team users. While OpenAI acknowledges limitations—such as occasional hallucinations and challenges in source verification—its iterative deployment strategy and continuous refinement approach are promising.

My key takeaway: This LLM agent-based tool has the potential to save hours of manual research while delivering high-quality, well-documented outputs. Automating tasks that traditionally require expert-level investigation, it can complete complex research in 5–30 minutes (just 6 minutes for my task), with citations and structured reasoning.

I don’t see any other comments yet from people who have actually used it, but it’s only been a few hours.I’d love to hear how it’s performing for others. What use cases have you explored? How did it do?

(Note: This review is based on a single use case. I’ll provide further updates as I conduct broader testing.)

[1] https://www.dazzagreenwood.com/p/ueta-and-llm-agents-a-deep-...