hahaha, makes me sad and happy all at once...
HN user
silksowed
Flying drones with natural language input instead of remote controllers. Instead of having a human on the control sticks the entire time, what if they could describe the goal they are trying to achieve, and then the drone goes and flies according to the agreed intent? I know traditional drones already have autopilots, but they seem to be static pre-planned routes and could benefit from advancing the capability to be more dynamic and flexible. Eventually I want to combine a live camera stream to run a local vision model that would identify and notify images of interest the drone sees while flying. From there the operator can decide if they want to re-task the mission or adjust as needed. That is the general technology direction, but I hope to expand this into the wildland firefighting vertical. Instead of putting humans up in helicopters and fixed wing aircraft, I hope we can leverage drones instead. I understand not every use case will be able to replace with a drone, but enough of them could be to make it worthwhile in my opinion. Still early days of research and development, but super excited about the tech and combining AI + robots for the common good. Currently a huge fan of the startup Seneca, and hope to help expand the industry or join them.
1) I recently published my latest milestone here: jakedecamp.com
Computer-use is a big limitation that my 2015 Macbook Pro cannot handle. I find the Codex cli says it looks at the end output artifact but so often it fails to refine it into acceptable form. If it could use my computer screen and visual inputs for review, it might be able to actually design documents/powerpoints/etc. I'm juicing everything I can out of the 11 year old laptop and I'm honestly impressed at what it can still do.
Same here. I find the design, architecture, system design discussion to be better on Claude, but after Opus 4.6 I switched over to Codex for actual coding and love the results. I use both via the CLI and generally tell Claude to output the result of our decisions as a markdown that will be easy to read and implement by an agentic coding tool. Then I fire up Codex and read said markdown as the input of the session and way to build all the appropriate context needed. I see this as a way to step into letting the agents go run on their own and interact with each other, but I still like to steer so I put these manual steps in the flow. Letting the agents go off on their own and one shot big chunks is not reliable enough yet imo.
Well said. Deciding what to build is derived by experience in the context of your problem space, and that is human centric. Makes me encouraged that as AI tooling increases in adoption, more people will build things to solve their problems. Instead of large corporations hoarding talent and then building things for people, I could see more small business success where they sit closer to the very things that need solving.
Interesting. I've been building around that MCP abstraction and have had some early success flying in Cosys-Airsim (and Gazebo before that): https://github.com/jakedcmp/droneserver . I am starting to realize I need to break apart the MCP interface from all the other pieces of the stack for cleaner architecture but thats pending work. My flow goes like this: LLM -> MCP tools -> droneserver -> MAVSDK/MAVLink -> PX4/Ardupilot -> Cosys-Airsim Software in the Loop testing. What is novel for me is not having to learn how to fly a drone and bringing that capability into already existing technologies like PX4 autopilot. I have been attempting to code my own mission planning so I will check out QGroundControl as that might already be a solved problem and not worth building from scratch. I have also built the foundation of video streaming back from the drone so I can run video/image perception. Once I get perception working I am hoping I can build intent level autonomy where images are analyzed according to high level mission plan and potentially re-task the drone based on that. For example, the user issues simple command to fly around the property and scan for broken parts of the fence. During flight if an image of a broken fence is perceived, then the drone stops its flight and goes closer to capture additional imaging/video and document a gps location. Still hacking things together towards a real demo so the code probably wont port over well but idk. To anyone in this thread who wants to discuss further or collaborate let me know, it seems we are all working in a similar domain but from slightly different angles. Exciting area to build, I know there is big demand for solutions in this space over the next decade.
Very cool, I'm working on a similar project but using MCP as the flight input layer. Would love to dm and discuss more about how you built it or collaborate.
Is an FPV controller any different than a regular video game controller? I interested in how hard it is to fly a drone even in a sim environment.
Is crowd strike like a digital twin / virtual world sandbox for autonomous drones? Do you have any additional information I could check out? Been working on autonomous drone flights but eventually need a digital world to experiment in but have yet to reach that step. Debating working with Unreal Engine or NVIDIA omniverse but unsure what the right direction is.
Been experimenting with AI -> MCP driven drone orchestration/flight would love to learn more about what your building and compare. Thoughts?
very excited to play around. will be attempting to see if i can get character coherence between runs. the issue with the 8s limit is its hard to stitch them together if characters are not consistent. good for short form distribution but not youtube mini series or eventual movies. another comment about IP license is indeed an issue but its why i am looking towards classical works beyond their copyright dates. my goal is to eventually work from short form, to youtube to eventual short films. tools are limited in their current form but the future is promising if i get started now.
how were you able to tell this? still trying to understand what infra is better used for inference (say realtime image category matching) vs training (feeding a chatbot huge sums of data)
face ID scan, click to confirm more power...! very good idea that i think would do very well considering the psychology of how people spend money
add in an equivalent to https://www.glean.com/ to enterprise dropbox and you have a new AI product that actually solves large org problems
i really wonder how they are housing the desktop grade hardware. im so used to rack and stack servers (1U/2U), but do you really need that big of a chassis for a desktop cpu, a couple ram dimm's and some ssd? what are you're thoughts?
it might be pastrami on rye time!
very interesting space, application sent!
ahh makes sense. not sure how far in the coastal commission would have control, but even a 10 story building far enough in would drastically change the landscape of say the sunset district in San Francisco
ocean sky scrapers on the west side of the bay area? hard to even picture this
totally understand where you're coming from. lots of days i feel like my job is to get the right people involved in the problem, not solve the problem myself. best of luck to you, i might find myself moving a similar direction.
i'm heading down the PM path from a CS degree background, working at a tech company for ~2 years. sometimes i wish i had gone into dev work so i could get better satisfaction in the day to day. it also seems like PM hiring is very picky which worries me if i lose my job. anecdotally lots of career pages have lots of open dev positions but few PM roles. is this a factor in your thought process for switching?
does anyone know of anyone using graviton for compute instances? would like to gauge how the experience has been
for GPU programming what area's would you suggest to start with? should i just dive into CUDA?
had a footprint in radiation oncology, at least at the time i interned. lots of machines that go into hospitals that are used on you but you never see
so if we run the traffic through a VPN we can connect existing infrastructure to mobile network devices? i'm very intrigued by the intersection of mobile networks and the existing TCP/IP that runs most of the modern data center. curious how those two pieces will evolve and communicate together. as a side note, any edge device that is outside of wifi network range also seems to be isolated. until you bring that device back into range there is no way to transmit the data, leading to localized storage issues the way i see it.
related but not related, how does a security camera transmit data via TCP/IP without a mobile network connection? does it run off a local wifi network to send and receive data? is there an interface to translate data from a mobile network connection back to TCP/IP allowing you to access via the IP addr?
us baby boomers own one of the largest asset classes in the world. so much wealth tied up in ~70 million people. i wonder if any of the children of this class are stepping off the gas knowing the windfall that eventually is passed on to them.
this just reminds me how US centric some of us default too, myself included.
any advice? my manager left 9 months ago and they decided to never backfill so i own basically all of it. granted i don't perform at the level he did but my pay has been stagnant and its always awkward in interviews to talk about projects that are clearly above my job title. they've floated the idea of a title bump/increase pay but it just seems like a carrot to keep me churning.
never thought of this but glad you said it. one thing i also think about if i was manager was creating a resilient system of employment. what happens if x leaves tomorrow? is x too critical and knows too much domain knowledge? how am i as a manager going to share that knowledge across my team so if any one person leaves we at least have some way to continue on.