HN user

aspectrr

130 karma
Posts6
Comments32
View on HN
[dead] 2 months ago

Hey HN,

I've been working on a way for agents to query production systems to help me debug issues and close the loop on things I work on day to day. It works as a hook that rewrites ssh, awscli, gcloud, az, kubectl commands to verify they are read-only and safe. It also keeps track of sessions in files and when agents debug the same things it will give hints in the tool calls like

━━ Past Investigation (May 10, 87% similar) ━━ Root cause: php-fpm pool exhaustion causing nginx 502 Hosts involved: web1 Investigation path: web1: systemctl status nginx web1: journalctl -u nginx --no-pager -n 20 web1: systemctl status php-fpm Consider checking: systemctl status php-fpm

Tools with memory is an interesting idea as well but lmk what you think!

[Lily](https://github.com/aspectrr/lily) A CLI tool that can be installed to any coding agent via hook that gives read-only access to production systems (wraps ssh, kubectl, awscli, gcloud, az) so agents can investigate issues in production. Built it for myself and my team during initial investigations to save use a lot of time on figuring out issues but didn't want to have to babysit agents or just hope that "telling them they are in production" would prevent issues.

[clue.ssh](https://github.com/aspectrr/clue.ssh) A clue game over SSH based on the AI wave, where the goal is to find who stole the H100. Pretty fun and coding agents can play too.

[Chasing Losses](https://github.com/aspectrr/chasing_losses) I was interested in if LLMs chased losses when playing roulette, still investigating this but i've found that different models will bet different amounts at different frequencies even when prompted the same. Struggling on not wanting to guide them too much but also wanting to see how they react when put under pressure.

[dead] 3 months ago

Hey HN, I have seen many different ways of letting AI run bash commands on remote hosts but none of which fix the issues of: a. safety (read-only) b. not installing anything on the remote host

so this is my implementation of one that does.

It uses seven layers of verification on the client and reconstructs the commands with safe quoting to prevent unsafe chars or other attack vectors. Check out: https://github.com/aspectrr/lily?tab=security-ov-file

Looking forward to your thoughts!

[dead] 3 months ago

Hey HN, I have seen many different ways of letting AI run bash commands on remote hosts but none of which fix the issues of:

a. safety (read-only) b. not installing anything on the remote host

so this is my implementation of one that does.

It uses seven layers of verification on the client and reconstructs the commands with safe quoting to prevent unsafe chars or other attack vectors. Check out: https://github.com/aspectrr/lily?tab=security-ov-file

Looking forward to your thoughts!

Hi HN,

My name is Collin, I've been working on automating my job and open-sourcing the results. I work as an ELK engineer and don't like so i started building this on my own time to find out if this was something that could be handled by agents and found success! The coolest part of which is built with sandboxes that have data stubs (kafka, s3, api) so the agent can model data pipelines in a full feedback loop without touching a cluster. Because of this I am working on an Elasticsearch consultancy comprised of me and a swarm of these agents working to build client projects.

Let me know if you have any questions!

[dead] 5 months ago

In the land of infrastructure, servers are sacred. Humans are barely allowed to SSH into these servers, and LLMs are not even in the picture. This is for good reason, one misspelled command and production is down. This is the reality that I saw working in infrastructure. However, I believe that the jump that Claude Code gave software engineers will happen to sys-admins, platform engineers, and dev ops people alike. I wanted to let LLMs onto these servers, and let them do my boring debugging work, safely. So that's what I built with Fluid.

A safe, auditable way to let LLMs debug and manage Linux Servers. Redact secrets, IP addresses, and keys from LLM inputs, have custom allowlists without completely hindering the LLMs performance, and audit logs. And once you are ready, give the LLMs sandboxes of your Linux Servers, allowing them to fix issues all on their own, safely.

Give it a shot and lmk what you think!

[dead] 5 months ago

In the land of infrastructure, servers are sacred. Humans are barely allowed to ssh into these servers, and LLMs are not even in the picture.

This is for good reason, one misspelled command and production is down. This is the reality that I saw working in infrastructure. However, I believe that the jump that Claude Code gave software engineers will happen to sys-admins, platform engineers, and dev ops people alike. I wanted to let LLMs onto these servers, and let them do my boring debugging work, safely. So that's what I built with Fluid.

A safe, auditable way to let LLMs debug and manage Linux Servers. Redact secrets, IP addresses, and keys from LLM inputs, have custom allowlists without completely hindering the LLMs performance, and audit logs. And once you are ready, give the LLMs sandboxes of your Linux Servers, allowing them to fix issues all on their own, safely.

Give it a shot and lmk what you think!

[dead] 5 months ago

Hey HN,

Collin back again, this time explaining how the new read-only mode works in fluid.sh, letting AI work on-prem.

If you have any questions or comments, I am happy to discuss more!

Yo, fluid is built with on-prem in mind, specifically VMs. This is my initial use case for it. I am currently working on a remote version of fluid, where instead of CLI tool, it would be more of a codex/claude code app with a UI where you can install a server and then command hundreds of agents at once to work on infrastructure. Is this what you had in mind?

For example, if you had an on-prem footprint with thousands of VMs, a production cloned sandbox would be a clone of a VM to let AI safely make changes, install packages, etc.

Yeah, working on the landing page. Feel free to ask any other questions!

Hey no problem! I'll work on the demo more. I discuss this in my comment here: https://news.ycombinator.com/reply?id=46889704&goto=item%3Fi...

and on the website: https://fluid.sh

But fluid lets AI investigate, explore, run commands, and edit files in a production-cloned sandbox. LLMs are great at writing IaC, but the LLMs won't get the right context from just generating an Ansible Playbook. They need a place to run commands safely and test changes before writing the IaC. Much like a human, hence the sandbox.

Hey! Yes I updated the website with some more of my comments. - RO mode would be a good idea - Agreed on explaining destructive actions. The only (possibly) destructive action is creating the sanbox on the host, but that asks the user's permission if the host doesn't have enough resources. Right now it supports VMs with KVM. It will not let you create a sandbox if the host doesn't have enough ram or cpus.

- The kubernetes example is exactly what this is built for, giving AI access is dangerous but there is always a chance of it messing something. Thanks for the comment!

Hey, thanks for the comment. I answer this question in more depth on the website https://fluid.sh or this comment: https://news.ycombinator.com/reply?id=46889704&goto=item%3Fi...

This lets AI work on cloned production sandboxes vs running on production instances. Yes you can sandbox Claude Code on a production box, but it cannot test changes like it would for production-breaking changes. Sandboxes give AI this flexibility allowing it to safely test changes and reproduce things via IaC like Ansible playbooks.

Hey, I get it. I don't want LLMs on prod at all. I made this to let agents connect to production cloned sandboxes, not production itself. I hope this helps your concerns, but I understand either way. Lmk with any other questions.

Thanks! Kubernetes is the next infrastructure primitive that I want to support but I'm glad you like. If you have any questions or ideas, lmk!

This allows the agent to make any changes in a production clone vs agents running on a production VM. For example, you wouldn't want claude editing crucial config on the chance it brings everything down vs letting it do in a cloned environment where it can test whatever.