HN user

ohans

21 karma

ohansemmanuel.com

Posts7
Comments22
View on HN

Author here.

AI agents have dominated Twitter/X the past few days: Claude Code, Remotion, Cursor, etc. But most developers are running these locally. Terminal agents on your machine.

That's useful, but it's a different game when you're building AI apps for users. The moment your agent needs to actually execute something e.g., run a Python script, process a PDF, generate files, you're suddenly building sandboxing infrastructure, managing VMs, handling file storage.

It's a massive distraction from your actual product. We built Bluebag to solve this for Vercel AI SDK users. Two lines of code, and your agent gets access to Skills that execute in managed sandboxes.

No Docker orchestration. No K8s. Dependencies, file handling, signed download URLs, all handled.

The post walks through some of the architecture (progressive skill loading, auto-provisioned VMs, multi-tenant isolation) and shows concrete integration examples.

Happy to answer questions about the approach or trade-offs we made.

Cheers

TIL: you could add a ".diff" to a PR URL. Thanks!

As for PR reviews, assuming you've got linting and static analysis out the way, you'd need to enter a sufficiently reasonable prompt to truly catch problems or surface reviews that match your standard and not generic AI comments.

My company uses some automatic AI PR review bots, and they annoy me more than they help. Lots of useless comments

Logging sucks 7 months ago

This was a brilliant write up, and loved the interactivity.

I do think "logs are broken" is a bit overstated. The real problem is unstructured events + weak conventions + poor correlation.

Brilliant write up regardless

Yes!! The runtimes are ephemeral VMs/containers with no network access (exposed)

On OS, the core of the solution is not currently open source; it’s still changing a lot, and I don’t want to publish an API or SDK surface that I’ll immediately have to update.

But I plan to open-source the CLI package and SDKs shortly

Hi HN,

When Anthropic published their Skills system (https://www.anthropic.com/news/skills), the idea clicked for me immediately: take a general-purpose agent and turn it into a specialized one with procedural knowledge that no model can fully memorize.

In my own projects I wasn’t using Claude (most of my workloads were on Gemini 2.5 Flash, mostly cos it was affordable and got the job done), but I still wanted that architecture: a way to define Skills once and use them with whatever LLM made sense for a given use case.

So over the past few weeks I put together a solution that does roughly that. Right now it supports:

- Bundling metadata, instructions, reference files, and optional scripts into a Skill - Running scripts in Python or JS runtimes (with automatic package installation) - A simple files API so the LLM can create files, reference them, mint temporary download links, and let me upload docs for analysis - A CLI to manage skills locally (push/pull), a Typescript SDK and a web app to manage API keys, PATs, playground etc.

There’s a playground at http://www.bluebag.ai/playground with example Skills (mostly adapted from Anthropic’s public Skills repo at https://github.com/anthropics/skills). On the right-hand side you can see how different models progressively load files and metadata, so you can inspect how selection and loading behave across models.

There are still some open questions I’m thinking about, especially around VM reuse and isolation at scale, and how to handle large Skill libraries over time (cold starts with very large package sets and 15+ Skills are slow).

But it’s been useful enough in my own work that I wanted to share it and get feedback. I’d be interested in:

- obvious failure modes I’m missing - prior art I should be looking at (e.g., agent frameworks)

Happy to answer any questions or dig into implementation details if that’s useful.

Cheers

You make some solid points!

I don't know what you mean by "true alternative to Amazon SES.

e.g., the GMail API is restrictive (https://support.google.com/a/answer/166852?hl=en) for bulk sends. That's not the intended use case.

Google partners with Sendgrid as well for customers who need that kind of solution.

Do you mean via the marketplace integration?

I could ask why Amazon doesn't offer a true alternative to GMail and Outlook.

That's fair. I reckon different models like you mentioned

Despite how crucial it is, it's hard to find a service that has all the features you need for a successful product.

That's why we built Stack.

I’m not convinced this is a strong USP. One could make a decent argument “Stack” doesn’t have “all” the features (yet) - I’ve seen the roadmap.

Firebase, Superbase etc. arguably have more features.

In a nutshell, I think it might be better to have that paragraph really drive home the USP of the product e.g., open source accountability, fairer pricing model etc.

Those are strong tells, instead of the “all the features” narrative. That’s a hard battle to win :)

Regardless, awesome work! And congrats on the launch