HN user

awwaiid

493 karma

meet.hn/city/us-Washington

Socials: - github.com/awwaiid

---

Posts5
Comments233
View on HN

Was there pressure to do this, or freedom to do this? If I had an unlimited token budget I'd probably try all sorts of crazy things. Also you (one) can read the tests and see that they weren't modified to forcibly pass.

Maybe he didn't think it would work. Maybe even if it does "work" they'll keep the zig version anyway. Maybe further study is needed beyond existing compiling/test-suite. Intentions and perspectives change over time, even only a few days, without dishonesty.

I'm guessing that if I said it ... that we have no intention of re-writing in rust ... that what I mean is "we have no intention of spending the extreme cost it would take to rewrite". When I discover the cost model is completely different that changes things.

Not AI generated software -- DYNAMICALLY generated software, like at run time and ongoing. Even in-app directed by the user. This is not a thing that existed before, a degree of customizability well beyond letting the user pick a color scheme or from one of a few layouts or default start screens.

I don't know how good of an idea it would be, product-wise, to give programming level flexibility. I am reminded of greasemonkey scripts, but written in english maybe. Maybe it could be awesome. But Apple is saying "nope. Not interested in exploring this with you. BYE"

Claude Opus 4.7 3 months ago

It's also difficult to recognize that when it got it right THAT might have been the lucky week.

Maybe another idea, no idea if this is a thing, you could pick your block-of-layers size (say... 6) and then during training swap those around every now and then at random. Maybe that would force the common api between blocks, specializaton of the blocks, and then post training analyze what each block is doing (maybe by deleting it while running benchmarks).

Microgpt 5 months ago

Where is this 1000 lines of C coming from? This is python.

I've been building a skill to help run manual tests on an app. So I go through and interactively steer toward a useful validation of a particular PR, navigating specifics of the app and what I care about and what I don't. Then in the end I have it build a skill that would have skipped backtracking and retries and the steering I did.

Then I do it again from scratch; this time it takes less steering. I have it update the skill further.

I've been doing this on a few different tests and building a skill which is taking less and steering to do app-specific and team-specific manual testing faster and faster. The first times through it took longer than manually testing the feature. While I've only started doing this recently, it is now taking less time than I would take, and posting screenshots of the results and testing steps in the PR for dev review. Ongoing exploration!

It was downplayed at every other opportunity including the conclusion, emphasizing instead the lone team of two heros. A few shout outs here and there, but the theme was clear.

Fine... I guess it could be happy ChatGPT day then? Putting a hosted version of it out on the internet and letting people use it is very clearly a very small step, it even surprised OpenAI that it caught so many people's attention. But it did. It is an interesting milestone that shows what happens when you connect people with a new technology, a new concept, and see what they do with it.

ChatGPT Atlas 9 months ago

I asked Atlas about this, and it indirectly pointed out that atlas://credits is a thing. Not linked to anywhere that I could find though.