Can you give an example you think will get cheaper compared to the average income in this future?
HN user
nevertoolate
Can you explain what you are working on?
I’ve stopped using llms to generate architecture, which i design and write myself and let the machine pattern match the gaps. I also use it to review issues which I lot of the times push back against.
I’m working on a stateful application sitting on top of a data warehouse and have to implement a stream of messy half defined feature requests and navigate on top of an ever changing infrastructure layer. LLMs rarely get the infra layer even if it is written as code and have hard time grasping how to deal with tech debt, when and how to re-architecture parts of the stack or even implement stuff based on a detailed openspec design.
Depends on the goal. If staying at a certain number is the goal yes, climbing two weeks back up will possibly take more and more time. If staying energetic and healthy is the goal the number becomes only one of the moving benchmarks.
With that said I think doing some kind of workout even on vacation is important.
I think tantrum comes when they are tired / disconnected from adult monologue. I have almost zero issues when I talk about interesting stuff instead of engaging in debate:
we will pick up that book about the monster, it is really scary (slippers on already) and we will sit on the sofa (already carrying the child). Are you cold? Let’s find that pink sweater…
So you think that having your words rewritten by an llm is somehow a more powerful, faster, safer way to write down your ideas? What do you want to say with this exactly? How you separate the “what” to write from “how” to write it?
Why I engage with this comment? I truly believe that trusting your own knowledge and skills is the way forward. If you don’t have the skill? Build it. Don’t outsource thinking and learning.
def my_swe_percentile(best_agent_swe_percentile):
return min(100, best_agent_swe_percentile * 1.25)I think it is fine to create the scripts with the cloud based llm but it is definitely not a fable / opus level thing, and running the bisect loop itself has nothing to do with an agent, it is a simple shell script.
I was trying to find the root cause of a crash in a Python module which left no errors in the log or console. Fable wrote a test harness that simulated clicks in the UI, then bisected my code until it found the point where it started crashing
Does this need an agent though is my question? Maybe generating a test case and a loop doing git bisect but why on earth would we want to run it through the internet and gpus and whatnot when it can be run on a single core celeron.
So you believe that your work will be done by AI and you will enjoy life more? This is not a loaded question, just trying to understand what your future ideal day / week would look like as an "ai optimist".
What is an ai enterprise tool?
I also notice these things. Otoh i spend definitely less than 50% of my time typing in code so it is impossible that it gives more than 2x speedup. And sometimes i lose time babysitting and rewriting stuff so all in all it is kinda no productivity gain.
If it was just programming being automated, then whatever.
There is nothing on horizon which automates a programmer’s work. Typing in code is faster now, and some things “only need pointing out” like an existence of a “bug” which an llm + harness might be able to mitigate. Automated tests might capture regressions and possibly written by llm + harness. If you replicate this in other professions what will you get?
My time as an experienced software engineer is worth a lot of money - a whole lot more than $12,000 for the past six months
From this I assume you think that what the llm has generated is as valuable as your own work generally is. How do you even calculate this?
Yes
So you don’t understand what you generate with ai and think that it will be a solution for a problem you can only solve using sql.
it’s cold -> turn on the heater
I’d never just turn on the heater silently if someone said this to me. I think it means something else.
What do you base this on? For me it is almost impossible to guess what fits into the context of an llm. Sometimes trivial tasks fail, sometimes quite complex things get one shotted.
I think it is great!
The issue is that validation needs presence and it is the limiting factor - common knowledge, but is part of the “physics”. Also maintenance gets really tricky if the codebase has warts in it - which it will have. I get much more easy to understand architecture out of an LLM driven code generation process if I follow it and course correct / update the spec process based on learnings.
Example: yesterday I’ve introduced a batch job and realized during the implementation phase that some refactoring is needed so the error boundary can be reused in the batch application from the main backend. This was unplanned and definitely not a functional requirement - could be documented as non-functional. There was a gap between the agent’s knowledge and mine even though the error handling pattern is well documented in the repository. Of course this can be documented better next time if we update the process of openspec writing but having these gaps is inevitable unless formal and half-formal definitions are introduced - but still there needs to be someone with “fresh eyes” in the loop.
I understand that you are serious. I am also serious here.
Have you built anything purely with LLM which is novel and is used by people who expect that their data is managed securely, and the application is well maintained so they can trust it?
I have been writing specifications, rfcs, adrs, conducting architecture reviews, code reviews and what not for quite a bit of time now. Also I’ve driven cross organisational product initiatives etc. I’m experimenting with openspec with my team now on a brownfield project and have some good results.
Having said all that I seriously doubt that if you treat the english language spec and your pm oversight as the sole QA pillars of a stochastic model transformer you are making a mistake.
Poe’s law is strong with this one
I'd suggest you to work on your general mood - drugs can help, but nature is also wonderful.
I think I have a relatively good life, but I still have hard times. I had circa 6 months long depression streak after my child was born (I'm male).
For me the best mood fixer is a walk still. Super small commitment, great with a dog too. For a weekend the best is a longer hike. I practice yoga and train my body - great mood boosters. I've trained my body to be able to sit comfortably on the ground so I can work from anywhere - sunshine in park hellooo.
Hope you find your rhythm soon!
So you had generated 2000 lines in 30 minutes and ran out of tokens? What was your prompt?
I’d use a fast model to create a minimal scaffold like gemini fast.
I’d create strict specs using a separate codex or claude subscription to have a generous remaining coding window and would start implementation + some high level tests feature by feature. Running out in 60 minutes is harder if you validate work. Running out in two hours for me is also hard as I keep breaks. With two subs you should be fine for a solid workday of well designed and reviewed system. If you use coderabbit or a separate review tool and feed back the reviews it is again something which doesn’t burn tokens so fast unless fully autonomous.
I wouldn’t be surprised and most likely would go for a walk :)
I agree - we should use the tools. But we should be mindful about how humans actually learn.
Some improvement ideas:
A prototype can help in the “Better communicate the idea/feature” part but it is even better if you let engineers do this as learning by doing is better than just being shown the result.
Vibe coding doesn’t help in “Understand the systems” - on the contrary, this is already a well known fact that vibecoding has negative effect in understanding the underlying system. It should be hardboiled documentation reading, trial and error which helps, otherwise you get only the illusion of competence.
Sounds like you are in a wonderful relationship, I’m glad!
My summary: openclaw is a 5/5 security risk, if you have a perfectly audited nanoclaw or whatever it is 4/5 still. If it runs with human-in-the-loop it is much better, but the value is quickly diminishing. I think llms are not bad at helping to spec down human language and possibly doing great also in creating guardrails via tests, but i’d prefer something stable over llms running in “creative mode” or “claw” mode.
So why are you stuck with ink/react stack?
Sorry but I don't understand why you ask this question, can you explain your train of thought?
Taiwan is a different story. There are quite detailed war simulations built for defending the country. I guess you might mean that russia is one of the “3 big world” powers and their move is the special operations to capture kiev. I stop here
- how to prove that humans can argue endlessly like an llm?
- ragebait them by saying AIs don’t think
- …