HN user

tekacs

4,675 karma

Amar Sood

Building pervasive.app

hn @ <username> . com

On GitHub, Twitter, etc. as tekacs.

If you've replied to me and I've not got back to you... I'm probably busy. I'm trying to change that.

ac5ecc4e73f4852ebd856bd3959cc2b611802b6330356a61088dabc826d60c86

meet.hn/city/us-New York

Posts162
Comments782
View on HN
www.youtube.com 10d ago

I built my DREAM New York City apartment from SCRATCH in 93 days [video]

tekacs
1pts0
www.youtube.com 1mo ago

QuadRF, a modular 4x4 MIMO beamforming tile built with an open antenna arch

tekacs
4pts0
twitter.com 6mo ago

The next big thing in heart disease prevention is targeting lipoprotein(a)

tekacs
1pts0
twitter.com 7mo ago

Sanders: Pushing for a moratorium on AI data centers

tekacs
2pts0
developer.chrome.com 8mo ago

Changes to remote debugging switches to improve security

tekacs
3pts1
help.openai.com 8mo ago

ChatGPT: You can now interrupt long-running queries to refine what you're asking

tekacs
2pts0
github.com 1y ago

Show HN: Navigator Mode (Like Claude Code) for Aider

tekacs
16pts5
abcnews.go.com 1y ago

Trump says US will 'take over' Gaza

tekacs
57pts36
www.youtube.com 1y ago

An Electronic Chessboard Without Turns

tekacs
1pts0
phys.org 1y ago

New discoveries about how mosquitoes mate may help the fight against malaria

tekacs
2pts0
twitter.com 2y ago

Phthalate Content in Common Foods

tekacs
6pts0
medium.com 2y ago

What would Secret look like in 2024? Introducing Speakeasy

tekacs
1pts0
linear.app 2y ago

Linear Asks – Turn Slack requests into actionable issues

tekacs
8pts0
www.32al.io 2y ago

Eloquent: Redesigned Text Editing for Android

tekacs
2pts0
blog.redplanetlabs.com 2y ago

We reduced the cost of building Mastodon at Twitter-scale by 100x

tekacs
954pts355
hariraghavan.com 4y ago

PCG: A Better Multiple to Value Companies

tekacs
2pts0
techcrunch.com 4y ago

Zoom announces Zoom Whiteboard, gesture recognition among several updates

tekacs
6pts0
www.abstractops.com 4y ago

AbstractOps: The Scaling System for Your Company

tekacs
20pts6
blog.airplane.dev 4y ago

How to gain conviction to work on a startup idea for 10 years

tekacs
2pts0
kevinkuchta.com 5y ago

Building a url-shortener with Lambda – JUST Lambda

tekacs
2pts0
aws.amazon.com 5y ago

Avoiding overload in dist. systems by putting the smaller service in control

tekacs
3pts0
tailwindcss-typography.netlify.app 5y ago

Tailwind CSS Typography

tekacs
1pts0
linc.sh 5y ago

Linc – The Perfect CI/CD Pipeline for Your Front End

tekacs
2pts0
www.nasa.gov 5y ago

NASA to Reexamine Nicknames for Cosmic Objects

tekacs
3pts1
oilprice.com 6y ago

Small Lab Makes Big Breakthrough in Nuclear Fusion Tech

tekacs
27pts27
docs.gitbook.com 6y ago

V2 Differences in GitBook

tekacs
2pts0
omz-software.com 6y ago

Keyboard – Utilities for the Pythonista Keyboard

tekacs
2pts0
amazingmarvin.com 6y ago

Marvin – Customizable Task Manager and Daily Planner

tekacs
55pts16
fberriman.com 6y ago

Fruit salad: a scrum estimation scale

tekacs
3pts0
setapp.com 6y ago

Setapp – The best apps for Mac in one suite

tekacs
1pts0

Given how small the counterexample is... this feels like a great example of where a lot of interesting results are going to be found: not because they were super difficult, but because intelligence didn't scale, and until computers could do this for us, the number of people who seriously poked at many such things was low.

I'm very excited for the impact of this effect in science and medicine and other disciplines too.

I have these plan files as well, but it depends on the scope and scale of the things you're executing on, I think. However much detail gets put into the plan, it still doesn't help if part of what the model needs to understand is the fine-grained / perfect detail of a large surface area.

When I've asked Codex agents about things that were in their context window, they've never – to my experience – been able to actually retrieve something from before compaction when using the proprietary compaction endpoint. Instead, they've had to consult their actual transcript.

So... at least as of a week ago or so, I don't believe so.

I know a lot of people like to say that compaction makes this moot, but the level of detail you lose across compaction is wildly too much for most things that I do, unfortunately.

Perhaps if your plans don't have as much detail, or if you're not, for example, having a discussion with a lot of nitty-gritty then it's fine?

The lack of long context is the main reason that I still end up using Anthropic.

The worst is when you need it to hold for example a number of papers in its head, or large and complex materials that it needs full resolution on and your context window ends up being perennially at 16%. You have about five minutes of conversation and it compacts and then you have to wait for it to read that again, get to 16%... and repeat.

372 was not perfect, but it was so much better and a godsend. It turned that 12 to 20% into more like 40%.

https://github.com/tekacs/fast-rm

I've overridden my rm with this, which I threw together for fast-deletes of things like Rust target/ directories, and after seeing the GPT horror story, I taught it to flatly reject deletions directly under `/` and under home directories, with a message printing the path that it's trying to delete.

Not exactly a perfect mitigation, but given that the stated risk was the model mistakenly using the wrong $HOME, it seems like a reasonable safety. I should probably make it use an even scarier rejection notice, though.

I also... have backups.

I came here to say that this is presumably ORE/OPE (order-revealing/preserving) encryption, not FHE, but...

It is both remarkable and depressing how _little_ information is given, and how buried it is on the CipherStash website... _any_ information on what their security and/or threat model is, what is actually stored, how encryption and search works, or any trade-offs involved.

Just to list a few pages that tell you next to nothing:

https://cipherstash.com/docs/stack/reference/what-is-ciphers...

https://cipherstash.com/docs/stack/cipherstash/encryption/se...

https://cipherstash.com/docs/stack/cipherstash/encryption

I eventually found:

https://cipherstash.com/docs/stack/reference/security-archit...

Which... it sounds like 'searched without being decrypted' means... it encrypts your query against their fast KMS and uses that to compare against indexes that were also encrypted with the same KMS? And ORE/OPE is an optional mode when you want range support.

It's not my patterns – as I said, this is from bulk tests to characterise the models, including their refusals – with very very different inputs, too. No matter the conversation tone, you get these.

You wouldn't have come across these in coding – these are more for 'things Anthropic's team have decided aren't acceptable enough for their taste'.

I love this, although I can't help but think that a lot of agents will - for better or for worse - send you a bunch of PII.

Labs are trying to make long-horizon work. Even if you're a coding agent, adding more and more surface area is distracting to that goal. There is reason that RL over long traces should, at least in principle, optimize for building in ways that help the result fit in the model's context window.

A meaningful risk of course is that the tools available to the model (ripgrep + fancier semantic approaches) allow it to do a good job of reasoning over things much larger than its context window, and so it doesn't pay the penalty sufficiently to fix it.

Absolutely this: and it needs to ideally become the kind of set of abstractions that mean that every new thing added uses less net-new surface area than it would without them.

I mean that at the bottom of the Tetris board, the lines need to vanish so that the Tetris board keeps moving downward and doesn't grow unbounded.

So is Spectral, which is mentioned in the headline of the article! As it says there:

SCALE delivers nearly a 6x performance boost on AMD GPUs compared to using HIPIFY to convert CUDA code to AMD’s own ROCm environment

... whilst also running CUDA.

I've said for a long time that composability in software is a bit like playing Tetris: the lines have to clear.

I feel like that gives an even more literal tower-rising metaphor, and that's what it feels like people using agents naively (and software engineers of lower skill or earlier-career), end up violating.

Agents are getting better at folding things into themselves, especially if you direct them to... but unfortunately I've found that the architectural instincts, even of Fable and 5.6 Sol, are still wildly behind what I reflexively achieve, say.

For sure there is an ability to have agents go back over work and try to fold it into better and better abstractions until it's sort of annealed into something good. I've done something similar on codebases that I have, but the 'high reaches' of architecture with great _prediction of how the software will evolve in the future_ in _subtle_ ways – those are, for now, out of reach of agents.

There is a part of me that wonders if it's partly just how much they can hold in their head right now, though. Even with the greatest articulation and high density of feeding them, the current setups don't allow them to hold a high-quality, sparse, 'zoomable' model of the world in their head that well yet, which we can do pretty well.

But the fact that I'm talking about it in terms of that kind of subtlety is itself promising, I guess?

I'm quite worried about the way that Anthropic in particular have trained their models to implement what they believe to be safety.

When the model has been trained not to do something [1], in my large-scale benches of such, it always says things in the spirit of:

- "... and that's a line I'd rather hold. Happy to <other things>"

- "I'm genuinely happy to <blah>, but I'm not comfortable with <blah>"

- "I don't want to keep going in <blah> direction"

etc.

Basically, they use very emotional and personal preference language.

It's as if they've weaponized the language of interpersonal comfort on behalf of their beliefs about what a model should or should not do. It's deeply uncomfortable and impolite for a human to ask a model to keep on doing something after it's expressed something this way, naturally. Even worse, it's all but guilt-tripping anyone who comes across it into the idea that they're doing something deeply wrong – exporting Anthropic's ideas about morality.

OpenAI, at least, have the decency to either just do a safety cutoff or keep it to a simple, "I can't do that."

[1]: I literally wrote 'when the model doesn't 'want' to do something' in my first edit of this comment, then caught myself. Case in point.

I code with AI all day, every day. But I do think that it's worth pointing to this issue (from March).

The author has said that they've redone it since, but the "from-scratch hand-built" framing specifically – for me – somewhat grates given the original heavy lifting from an existing AGPL codebase.

https://github.com/cesanta/elk/issues/75

I want to acknowledge that the original authors don't seem to have minded too much – per that thread – after older versions were dropped.

For context, the current code doesn't look like it is the same shape, the same structure, etc., etc. – it _has_ been rewritten since (the 'since Feb' rewrite mentioned adjacent is related to this, AFAICT).

To the author: I absolutely love what you're doing overall. Keep going! Just be careful, folks.

GPT-5.6 13 days ago

Unfortunately, I'm finding that in long-form agentic use, when I'm trying to use Sol, I keep tripping guardrails – moreso than even Fable, somehow.

I don't know exactly what part of my codebase is triggering it, so I'm going to have to keep poking, but apparently the guardrails are not that gentle despite the phrasing. :(

ChatGPT Work 13 days ago

Anthropic just changed their web interface yesterday to have Chat versus Cowork as well, and every time I look at it, I'm so confused. I'm still so unclear when I'm supposed to use one or the other or the other.

Now the 'ChatGPT desktop app' (the Codex app, renamed) also has the split between work and code, and as far as I can tell, all it does is change which plugins are loaded by default to include Office ones when you put it in work mode. Perhaps it also changes the system prompt slightly?

I feel like a core difference is that the AI implementor can get cheaper/faster (and indeed _uniformly_ better), whereas it would be very difficult for the same humans to do so.

Even if this is not the right answer today, it can at the very least serve as a herald of a possible future, no?

https://developers.openai.com/api/docs/models/gpt-5.5-pro

GPT-5.5 Pro does not offer a cached input discount.

I think this tells you in one line. It's basically set up for one-shot inference right now, by the looks of things. If you use this in a harness, it would almost immediately fall apart on cost. Not to say that they couldn't make it work, just saying that at least as it's delivered currently, they haven't done so. On the web, there might be doing something to get the equivalent of that behavior internally, such as keeping the session truly alive on GPUs rather than using their external-facing cache-style approach.

CursorBench 3.1 20 days ago

I definitely use GPT-5.5 as a counterpart to validate these exact sorts of things in Anthropic models' implementations, in the (now-rarer) cases where I allow Anthropic's models _to_ implement.

And yeah, it's a bit depressing to think that 5.6 might be similarly nerfed. Less secure software for us all, I guess... except BigCorps. :(