HN user

westoncb

4,113 karma

Software engineer (usually early/founding stage) with a love for deep HCI and system design problems. Have been deep into this generation of AI, both in terms of designing/building systems on LLMs and in figuring out the highest leverage ways of using them for software dev.

http://symbolflux.com

https://twitter.com/Westoncb

westoncb@gmail.com

Looking for work!

Posts51
Comments1,297
View on HN
westoncb.github.io 1y ago

Show HN: Laser-Tracer, programmable virtual volumetric vector display

westoncb
1pts1
github.com 1y ago

Show HN: Real-time nonlinear optics simulation (JS/GLSL)

westoncb
46pts12
twitter.com 2y ago

Languages for LLMs to Think In

westoncb
2pts0
github.com 3y ago

Show HN: Cross-platform Desktop front-end for Stable Diffusion and others (MIT)

westoncb
1pts0
github.com 3y ago

Show HN: GenerationQ – Open Source Desktop GUI for Stable Diffusion and others

westoncb
8pts1
news.ycombinator.com 4y ago

Show HN: Simple Lambda Calculus Visualizer

westoncb
3pts0
news.ycombinator.com 6y ago

Ask HN: Are there reasons to avoid a Sublime Text-like business model?

westoncb
2pts11
symbolflux.com 6y ago

Show HN: Lucidity – an interactive program-state visualizer

westoncb
184pts35
neon-bindings.com 6y ago

Neon: Rust Bindings for Node.js

westoncb
2pts0
github.com 6y ago

PlayCanvas: Fast and lightweight WebGL game engine

westoncb
3pts0
80.lv 7y ago

Cascadeur: physics-based character animation tool

westoncb
1pts0
github.com 7y ago

Show HN: Minimal game with procedural graphics in JavaScript/GLSL

westoncb
34pts11
github.com 7y ago

Show HN: Minimal game with 100% procedural graphics in JS/GLSL

westoncb
6pts3
www.quora.com 8y ago

Why is linear motion relative but rotation absolute?

westoncb
3pts0
news.ycombinator.com 8y ago

Ask HN: What's the best way of learning calculus, if you already know pure math

westoncb
2pts4
blogs.msdn.microsoft.com 8y ago

A History of the Windows Command-Line

westoncb
8pts0
symbolflux.com 8y ago

Show HN: Interactive Conway's Game of Life with good graphics and writeup/source

westoncb
5pts1
news.ycombinator.com 8y ago

Ask HN: What's your favorite way of getting a web app up quickly in 2018?

westoncb
542pts558
news.ycombinator.com 8y ago

Ask HN: Is there a clear, mature discussion of OOP/functional tradeoffs?

westoncb
4pts4
westoncb.blogspot.com 8y ago

Heuristics for Choosing 'Important' Books to Read

westoncb
1pts0
www.nytimes.com 8y ago

Astronaut Scott Kelly: How Tom Wolfe Changed My Life

westoncb
45pts4
westoncb.blogspot.com 8y ago

Thoughts on How to Find Alternate Algebra-Like Systems

westoncb
3pts0
westoncb.blogspot.com 8y ago

Heuristics for Deciding a Book's Reading-Worthiness

westoncb
1pts0
news.ycombinator.com 8y ago

Ask HN: Can you recommend resources for learning JavaScript app architecture?

westoncb
2pts0
symbolflux.com 8y ago

Show HN: Interactive Conway's Game of Life with good graphics (MIT License)

westoncb
5pts1
news.ycombinator.com 8y ago

Ask HN: Would Feynman be nearly as well known if it weren't for his personality?

westoncb
1pts1
westoncb.blogspot.com 8y ago

Thoughts on how to find alternate algebra-like systems

westoncb
3pts1
news.ycombinator.com 8y ago

Ask HN: Is there a good alternative to craigslist for finding rooms/housing?

westoncb
1pts0
www.sciencedaily.com 8y ago

Empathy represses analytic thought, and vice versa (2012)

westoncb
1pts0
westoncb.blogspot.com 8y ago

Thoughts on how to find alternate algebra-like systems

westoncb
8pts0

I've been building this "general problem solver" (will likely focus on math problems first) that uses a special kind of orchestrator to direct/structure the problem solving approach in accordance with how many 'rounds' remain and other aspects of problem context. It does this largely by influencing the behavior of specialists.

Just posted a first early demo and sample orchestrator system prompt yesterday: https://x.com/Westoncb/status/2053429329233895857

You initialize the system with an objective and a number of rounds to run for, and it loads the current config (orchestrator + specialist prompts and LLM configs) and begins working on it. You can manually step one round at a time or just let it run.

Rather than accumulating a single long work log/context, at each round specialists apply patches to a number of named 'artifacts' with different roles (e.g. uncertainties, dead ends, findings), which are injected into prompts during subsequent rounds.

The engine is written in rust and there's a web UI (and CLI). You can use the built in config editor to define specialists (and their prompts), what the artifact set is, orchestrator prompting etc.

Taking on a 'slow' software project with the kind of attention to quality (inside and out) that I had pre-AI. It's a tool I'll use myself, LLM-related, but not any kind of radical idea; it's main value is in careful UX design/efficiency, engineering quality, and aesthetics.

I've been shooting for the moon with one experimental idea after another (like many others) testing out LLM capabilities as they develop, for at least 2yrs now.

I'm still very excited about how these new tools are changing the nature of software development work, but it's easy to get into this frenetic mode with it, and I think the antidote is along the lines of 'slowing down'.

but you could give me two black boxes that act the same externally, one written as a single line , single character variables, etc. etc. etc. and another written to be readable, and I wouldn't care so long as I wasn't expected to maintain it.

The reality of software products is that they are in nearly in all cases developed/maintained over time, though--and whenever that's the case, the black box metaphor fails. It's an idealization that only works for single moments of time, and yet software development typically extends through the entire period during which a product has users.

I read OPs "good code" to mean "highly aesthetic code" (well laid out, good abstractions, good comments, etc. etc.)

The above is also why these properties you've mentioned shouldn't be considered aesthetic only: the software's likelihood of having tractable bugs, manageable performance concerns, or to adapt quickly to the demands of its users and the changing ecosystem it's embedded in are all affected by matters of abstraction selection, code organization, and documentation.

Interesting that compaction is done using an encrypted message that "preserves the model's latent understanding of the original conversation":

Since then, the Responses API has evolved to support a special /responses/compact endpoint (opens in a new window) that performs compaction more efficiently. It returns a list of items (opens in a new window) that can be used in place of the previous input to continue the conversation while freeing up the context window. This list includes a special type=compaction item with an opaque encrypted_content item that preserves the model’s latent understanding of the original conversation. Now, Codex automatically uses this endpoint to compact the conversation when the auto_compact_limit (opens in a new window) is exceeded.

That depends on the content of the SVGs.. Of course you can write a script to do a very literally kind of conversion of regardless, but in practice a lot of interpretation would be required, and could be done by an LLM. Simple case is an SVG that's a static presentation of a button; the intended React component could handle hover and click states and change the cursor appropriately and set aria label etc. For anything but trivial cases a script isn't going to get you far.

That's about how it came across for me as well: ignoring my actual content and joking about generalizations related to key words.

Project is cool overall, love the xkcd-like comic idea—but prompting and/or model-selection could use some work. I'd like to take a crack at tuning it myself :)

It sounds more like you just made an overly simplistic interpretation of their statement, "everything works like I think it should," since it's clear from their post that they recognize the difference between some basic level of "working" and a well-engineered system.

Hopefully you aren't discouraged by this, observationist, pretty clear hansmayer is just taking potshots. Your first paragraph could very well have been written by a professional SWE who understood what level of robustness was required given the constraints of the specific scenario in which the software was being developed.

Ah interesting, I missed that possibility. Digging a little more though my understanding is that what's universal is a shared basis in weight space, and particular models of same architecture can express their specific weights via coefficients in a lower-dimensional subspace using that universal basis (so we get weight compression, simplified param search). But it also sounds like to what extent there will be gains during inference is in the air?

Key point being: the parameters might be picked off a lower dimensional manifold (in weight space), but this doesn't imply that lower-rank activation space operators will be found. So translation to inference-time isn't clear.

So, they found an underlying commonality among the post-training structures in 50 LLaMA3-8B models, 177 GPT-2 models, and 8 Flan-T5 models; and, they demonstrated that the commonality could in every case be substituted for those in the original models with no loss of function; and noted that they seem to be the first to discover this.

Could someone clarify what this means in practice? If there is a 'commonality' why would substituting it do anything? Like if there's some subset of weights X found in all these models, how would substituting X with X be useful?

I see how this could be useful in principle (and obviously it's very interesting), but not clear on how it works in practice. Could you e.g. train new models with that weight subset initialized to this universal set? And how 'universal' is it? Just for like like models of certain sizes and architectures, or in some way more durable than that?

I think the idea is like: it took extra work 'cause Rust makes you be so explicit about allocations and types, but it's also probably faster/more reliable because that work was done.

Of course at the end of the day it's just marketing and doesn't necessarily mean anything. In my experience the average piece of Rust software does seem to be of higher quality though..

Doing math is not the same as calculating. LLMs can be very useful in doing math; for calculating they are the wrong tool (and even there they can be very useful, but you ask them to use calculating tools, not to do the calculations themselves—both Claude and ChatGPT are set up to do this).

If you're curious, check out how mathematicians like Robert Ghrist or Terence Tao are using LLMs for math research, both have written about it online repeatedly (along with an increasing number of other researchers).

Apart from assisting with research, their ability on e.g. math olympiad problems is periodically measured and objectively rapidly improving, so this isn't just a matter of opinion.

lol no problem. In reality though there's kind of a funny story behind it because I suspect the way I ended up using them so much is similar to how ChatGPT did. When I got into writing I studied grammar, then decided to read a bunch of classics and analyze their usage of punctuation in general until I had a good understanding of every bit of it. Then, in order to practice, I'd apply what I learned to anything I was writing at the time whether journal notes, conversations on AIM/IRC etc. That latter step meant I was translating a lot of casual/natural speech into a form that also had a high level of 'correctness'. And if you faithfully translate natural speech into 'correct'ly punctuated sentences, you end up using a lot of em dashes. Because ChatGPT/LLMs are tuned for natural/authentic style, as well as for a high degree of 'correctness,' you get today's state of affairs. Just a theory.

The difference is the incentive to improve, and actual present rate of improvement, for models like this is far higher than it is for jetpacks. (That and certain intrinsic features at least suggest the route to improvement is roughly "more of the same," vs "needs massive unknown breakthrough".)

lol yep, fully get that. And I mean I'm sure o4 will be great but the '-mini' variant is weaker. Some of it will come down to taste and what kind of thing you're working on too but personal preferences aside, from the heavy LLM users I talk to o3 and gemini 2.5 pro at the moment seem to be top if you're dialoging with them directly (vs using through an agent system).

I've seen that specific kind of role-playing glitch here and there with the o[X] models from openai. The models do kinda seem to just think of themselves as being developers with their own machines.. I think it usually just doesn't come up but can easily be tilted into it.

Gotcha. Yeah, give o3 a try. If you don't want to get a sub, you can use it over the api for pennies. They do have you do this biometric registration thing that's kind of annoying if you want to use over api though.

You can get the Google pro subscription (forget what they call it) that's ordinarily $20/mo for free right now (1 month free; can cancel whenever), which gives unlimited Gemini 2.5 Pro access.

There is a skill to it. You can get lucky as a beginner but if you want consistent success you gotta learn the ropes (strengths, weaknesses, failure modes etc).

A quick way of getting seriously improved results though: if you are literally using GPT-4 as you mention—that is an ancient model! Parent comment says GPT-4.1 (yes openai is unimaginably horrible at naming but that ".1" isn't a minor version increment). And even though 4.1 is far better, I would never use it for real work. Use the strongest models; if you want to stick with openai use o3 (it's now super cheapt too). Gemini 2.5 Pro is roughly equivalent to o3 for another option. IMO Claude models are stronger in agentic setting, but won't match o3 or gemini 2.5 pro for deep problem solving or nice, "thought out" code.

This sounds like a cool project and I may sign up. I think the idea of seeding collaborations with side projects vs something necessarily serious right out the gate is pretty solid/clever.

Edit: I do agree with others though that we should be able to see projects/profiles before registering.

Location: Tucson, AZ (US) Remote: Yes

Willing to relocate: Possibly!

Technologies: LLMs, Typescript, node, React, three.js, dabbled with Elixir and Rust, lots of Java a long time ago and a bit of Objective C/Swift

Portfolio: https://symbolflux.com

Github: https://github.com/westoncb

LinkedIn: https://www.linkedin.com/in/weston-beecroft-b4a98054

Email: westoncb@gmail.com

I'm an experienced generalist software engineer typically working with early-stage startups or with clients on relatively greenfield projects.

I have strong technical foundations, good product sense, and have gone deep with AI/LLM tech both in the sense of being able to use it effectively, and in exploring the design space for products/tools leveraging it.

I've done a lot of work in developer tools and data visualization of various kinds. Data-rich, non-traditional UIs with highly optimized UX, and rapid prototyping are my forte.

I've been basically on sabbatical for a while now, mostly taking the time to learn and build with AI. At some point in early 2024 I decided: okay, time to take this seriously, and have. After the long break, I'm ready to start something new!

SEEKING WORK | USA | Remote is fine (I'm also considering relocation to Chicago or NYC, maybe SF)

I'm a jack-of-all-trades software engineer who's done extensive "founding engineer" work in the context of both VC-backed startups and for smaller contract gigs over a period of about 10 years.

I have strong technical foundations, good product sense, and have gone deep with AI/LLM tech both in the sense of being able to use it effectively, and in exploring the design space for products/tools leveraging it.

Tools I use most frequently these days: Typescript, node, React, pnpm+vite. I've also done extensive work with three.js in the past. I used to write a lot of Java and a bit of Objective C, have dabbled in Rust and Elixir.

I've done a lot of work in developer tools and data visualization of various kinds. Data-rich, non-traditional UIs with highly optimized UX, and rapid prototyping are my forte.

Email: westoncb@gmail.com

Portfolio: https://symbolflux.com

Github: https://github.com/westoncb

LinkedIn: https://www.linkedin.com/in/weston-beecroft-b4a98054

and must say that I'm a bit confused by this presentation, and it's a bit unclear to me what it adds.

I think the disconnect might come from the fact that Karpathy is speaking as someone who's day-to-day computing work has already been radically transformed by this technology (and he interacts with a ton of other people for whom this is the case), so he's not trying to sell the possibility of it: that would be like trying to sell the possibility of an airplane for someone who's already just cruising around in one every day. Instead the mode of the presentation is more: well, here we are at the dawn of a new era of computing, it really happened. Now how can we relate this to the history of computing to anticipate where we're headed next?

...but sometimes an LLM becomes the operating system, sometimes it's the CPU, sometimes it's the mainframe from the 60s with time-sharing, a big fab complex, or even outright electricity itself?

He uses these analogies in clear and distinct ways to characterize separate facets of the technology. If you were unclear on the meanings of the separate analogies it seems like the talk may offer some value for you after all but you may be missing some prerequisites.

This demo app was in a presentable state for a demo after a day, and it took him a week to implement Googles OAuth2 stuff. Is that somehow exciting? What was that?

The point here was that he'd built the core of the app within a day without knowing the Swift language or ios app dev ecosystem by leveraging LLMs, but that part of the process remains old-fashioned and blocks people from leveraging LLMs as they can when writing code—and he goes on to show concretely how this could be improved.

I think the community here would find the methodology behind building this of interest. How all-encompassing "vibe coding" has come to be for "developing software with LLMs" is unfortunate imo. I have experimented with vibe-coding too, but it's a very different thing from the process used here. This tweet from Andrej Karpathy is the best description of how I approach development with LLMs: https://x.com/karpathy/status/1915581920022585597