HN user

azath92

252 karma

works at climatealigned. reach me: leobrowning92 on gmail

Posts2
Comments62
View on HN

Looking at the statistics that are used to measure growth and performance at a population level are such a great window into how they themselves are opinionated stances.

The two measures are both great measures of house ownership that measure different things. The article frames the second one as better, and if you are looking to see the ownership proportion of people (for example as they note to look at the numbers of young people who live with parents) thats great. If however you care about the housing development decisions that will incentivise owner occupancy you might care more about the standard owner occupancy per household.

I often see this with respect to GDP/GDP per capita vs mean income and other economic measures, but its great to see a non economic measure like this compared in this way.

That sounds like a tough realization to come to, and has been an opposite of my experiences in startups (as a non founder). Ive sought them out because the leadership cares so freaking much about what they are doing, and really genuinely wants to succeed. That can be intense, but at least you know they care. If you are early enough to have both equity and a solid chance at impacting the possibility of success you can have both mercenary and personal incentives aligned in a way that is impossible in a bigger company IMO.

I hope you give startups another go if that sounds good to you, and as you say, it sounds like you will be able to see the things you dont want at least. Good luck.

I can appreciate this sentiment. I am someone who favors a resign vote over a let it play in 8v8, and when i was thinking about why i think its that in a lobby based system like bar, where there are few games I value staying in a lobby rather than skipping around. Thus when a game feels lost to me, id rather play a new game (even to another loss) with a good lobby then play a slow inevitable losing struggle, or worse yet watch someone else play a losing struggle (even if they enjoy playing that loss out).

Something im trying to lean into in my games is to ask the players who are voting no what their win condition or what they think they can make a go of is.

I think what ive said above could make me sound like one of the folks who "want the tightest possible loop to winning and restarting", so to add some nuance there, I think playing a game where the conclusion is foregone lacks agency, and in a game where one team eventually loses of course there is a blurry area where the team agrees that they are in fact never going to win, and then another question on whether they lie to play regardless. Because of the variation there, i like the vote, and i am in favor of coms around why someone wants to resign or stay.

I actually think it is up to players in an open loby system to communicate their intentions and reasons, and while it is true that it means people who spam coms trying to push folks around are possible, i would echo ptaqs sentiment to use that report function! and on a softer level, i really like the mute. if someone is too much, mute them and continue to enjoy your game.

While i appreciate that this is a valid link. I would argue that its political stance is clear: its framing of this as "Biden’s border crisis", and its sourcing is poor: referencing fox news but no other news sources, and no links to primary documents or data, especially around the causality.

Additionally it doesn't appear to reference the actual HUD report from 2025 https://www.huduser.gov/portal//portal/sites/default/files/p...

This at least has AFAICT the 100% figure, but the causal link seems unsubstantiated. The exact, unreferenced quote from the document is "Immigration accounts for up to 100 percent of housing demand growth in some regions, and for two-thirds of rental demand growth nationwide. In California and New York, immigrants have accounted for 100 percent of all rental growth and over one-half of all growth in owner-occupied housing in recent years." p45 with no further citation to back that up. I know its not an academic paper, but some citation would be good to back this up.

Id be interested in more information on causal data, or at least a more nuanced stance on such a strong statement. Or a less politically charged comment.

Full disclosure, i (obviously from my comment) dont lean right, but one thing i appreciate about HN is a more nuanced comment, and i think this is a chance for that.

In a small team, or an aware team, where AI is being used all the time and we are figuring out the best way to do it, i often just preface my messages with

  - "from my ai to yours" where ive pointed my ai at some relevant context, and asked it to transform it for other ai context that a coworker needs
  - "my thoughts prettied by AI" where i just polished up my own words, often for outside coms, but indicating that i wrote the bones of it.
  - "i wrote this myself" in my case i tend to be very casual with my written coms, and ive been leaning into this in the past year rather than looking to correct it, as it gives the personal feel. but for cases where ive written more thoughtfully, i just flat out say that.
Now im not doing this rigerously, or obsessively, but i am finding it helps with exactly the kind of friction and erosion of trust that comes from reading things by ai as if i should treat it the same as a person and writing things as a person just to have it consumed and spat out again by an ai.

Helps my team is small. interested in how this could be translated to more widespread "company culture"

Id disagree with this analogy: "No carpenter is a specialist in drills." and i think its an interesting lens through which to look at the evolution of our tools.

I think there are trades where tool (or process if i may be allowed to extend the analogy) specialists exist and are highly valued. My dad is a plumber, so ill use that example but id trust similar is true for carpentry. there are specialists by task/output (new construction, repairs, boilers etc) but also tool specialist plumbers and companies for example drain clearing equipment or certain kinds of pipe for handling chemicals other than water are very specialised, and there are roles for them because the thing they enable, and the criticality of the task, and often the cost and complexity of using the tool are high enough to make specialisation valuable.

IMO software has, for the 10 years ive been working in it, been in an unusual position where the tools (languages, engineering practices, tech stacks) were super technical and involved, but also could be applied to a large number of problems. That is the perfect recipe for tool specialists: complex tool with high value and broad domain/problem space applicability.

Because of that tool specialisation, we've separated the application of the tool to a problem/domain from the tool use. reduction of complexity of applying these tools to many problems, means all domain specialists will use them, relying less on tool specialists.

imaging a mcguffin tool for attaching any two materials together, but which took a degree to figure out (loose hyperbole here), that sudenly you could use for 5 bucks and a quick glance at the first page of the manual. An industry that used to have lots of mcguffin engineers, would be mega disrupted, and you could argue that those tool specialists would have to identify more with what they were building than the mcguffin they were using.

the python cookbook is good. and fluent python is more from principles rather than application (obvs both python specific). I also like philosophy of software design. tiny little book that uses simple example (class that makes a text editor) to talk about complexity, not actually about making a text editor at all.

I made the same gut assumption, and it points to either poor writing, or deliberately misreading writing that they mix units like that in the same paragraph, where presumably the idea is that we get a feel for growth in both?

Its probably nitpick correct, because the 12GW is planned capacity, while the solar might be measured use? but simple assumptins or conversions, as another comment points out, get you comparable numbers. taking the title into account, the whole article is a little bit smoke and mirrors on clear communication, despite having plenty of numbers. Thats a shame because it sounds like even unvarnished its good results!

Id guess by your smile there is an element of humor in your response, so this isn't a rebuttal, but rather i identified a lot with your point, and I was thinking that this is such a human response to vulnerability.

If it was guaranteed that it will not be abused or that I would regret it, it would not _be_ vulnerable. Just like its not bravery if I am not afraid or I am assured of my safety. Such a paradox. Being vulnerable for me is acknowledging that it might have an increased probability of a more negative outcome, but still trying to be vulnerable because of the huge connection unlocks that (often) occur in my experience.

On balance intellectually i am coming to see the expected value from being vulnerable in communications is high, but my little lizard brain keeps saying to me "what if you get hurt though" and being closed off haha. its an exercise to shut it up.

I find the money stuff newsletter by Matt Levine (bloomberg) great for this, the link is behind a paywal, but the newsletter is free. strong rec. todays newseltter https://www.bloomberg.com/opinion/newsletters/2026-03-11/pri...

From that newseltter:

At the Financial Times, Jill Shah and Eric Platt report:

JPMorgan Chase ... informed private credit lenders that it had marked down the value of certain loans in their portfolios, which serve as the collateral the funds use to borrow from the bank, according to people familiar with the matter. >...

The loans that have been devalued are to software companies, which are seen as particularly vulnerable to the onset of AI. ...

From what i can tell the problem isn't that an individual who had cash to invest in a private (tech in this case) company goes down

the problem is that a company "private credit firms run retail-focused funds (“business development companies” or BDCs)" which took out a bunch of loans to invest in private tech companies is now having the underlying assets that they got those loans against (long term investments in private tech companies) valued lower.

the link im missing is what happens when people who also invested in BDCs want their money back, where their actual money is locked up in long term investments made to private tech companies, and their ability to get loans is now valued lower. I think this is called a "run" where if someone starts pulling money out, and ultimately you cant, then its a race to get your money out before others do, which applies to both the individuals and the institutional loans.

Note: my quotes are from the bloomberg newsletter i mention, which helped me, not the OP article. And i am writing as much to clarify my own thinking as from a place of understanding. I welcome clarification.

ok, even that "few thousand examples" heuristic is useful. the usecase would be to run this task over id say somewhere in the order of magnitude of 100k extractions in a run, batched not real time, and we'd be interested in (and already do) reruns regularly with minor tweaks to the extracted blob (1-10 simple fields, nothing complex).

My interest in fine tuning at all is based on an adjacent interest in self hosting small models, although i tested this on aws bedrock for ease of comparison, so my hope is that given we are self hosting, then fine tuning and hosting our tuned model shouldn't be terribly difficult, at least compared to managed finetuning solutions on cloud providers which im generally wary of. Happy for those assumptions to be challenged.

Only to prompt thought on this exact question, im interested in answers:

I just ran a benchmark against haiku of a very simple document classification task that at the moment we farm out to haiku in parallel. very naive same prompt system via same api AWS bedrock, and can see that the a few of the 4b models are pretty good match, and could be easily run locally or just for cheap via a hosted provider. The "how much data and how much improvement" is a question i dont have a good intuition for anymore. I dont even have an order of magnitude guess on those two axis.

Heres raw numbers to spark discussion:

| Model | DocType% | Year% | Subject% | In $/MTok |

|---------------|----------|-------|----------|-----------|

| llama-70b -----| 83 | 98 | 96 | $0.72 |

| gpt-oss-20b --| 83 | 97 | 92 | $0.07 |

| ministral-14b -| 84 | 100 | 90 | $0.20 |

| gemma-4b ----| 75 | 93 | 91 | $0.04 |

| glm-flash-30b -| 83 | 93 | 90 | $0.07 |

| llama-1b ------| 47 | 90 | 58 | $0.10 |

percents are doc type (categorical), year, and subject name match against haiku. just uses the first 4 pages.

in the old world where these were my own in house models, id be interested in seeing if i could uplift those nubmers with traingin, but i haven't done that with the new LLMs in a while. keen to get even a finger to the air if possible.

Can easily generate tens of thousands of examples.

Might try myself, but always keen for an opinion.

_edit for table formatting_

Why No AI Games? 5 months ago

I was aware of the patent, and agree i think its overly narrow and you could get around it easily. I think the reason we haven't seen it or something like it in another game (or i haven't but someone pleeeease id love to hear systems like it), is less because its not useful, or maybe its not useful as a plug and play because the only reason it works is because of the super exhaustive care taken on tuning its parameters and giving it enough variety to make it interesting to play.

Kinda like the dialogue/story paths in something like hades, where IIRC they made a whole system to manage it, but the reality is that system only matters when the tree is suuuuuuuuper complex. or maybe it was disco elysium, or both ...

Why No AI Games? 5 months ago

And to separate my thoughts from the info blob:

i think the culture war point is also super true of the game design industry, not just the consumers, where the already ultra competitive nature of the work means that the creatives and the industry as a whole have taken a veeeery strong stance against genai. Thats a reckon, and i dont know if its good or bad.

It does feel a little counter to the march of progress, but in a medium where high effort can be enjoyed by many, im personally cool with artisinal handmade games.

Why No AI Games? 5 months ago

Only because it is something i find fascinating: > There are those Orcs in that one Lord of the Rings game who hold grudges against you.

Is referring to the nemesis system in Middle-Earth: Shadow of Mordor and Shadow of War, and its an amazing set of interlocking procedural systems that do genuinely feel like its AI, but is really AI in the sense its always been used by games (the rules the games follow to govern NPCs+world) and not AI in the sense of modern LLMs or even other generative systems. This video is a great look at what it is and why its great IMO https://www.youtube.com/watch?v=Lm_AzK27mZY

I think a system like this could really work well with some modern LLM stuff, but it certainly feels magic without it.

Climatealigned | Onsite | London

We've spent the last three years making climate finance data at a fraction of the time and cost using AI. We started pre GPT and have continued to evolve and rebuild our stack with the times. These days we are fully agentic, powered by opus/haiku using python/js as best fits the job, but we ride the wave and aren't attached to the past. We are a small, focused, in-person, and technically capable team with deep industry connections.

We are looking for an early career builder to take ownership of end to end data creation, and who can stay on their toes with rapidly changing tech stack and ways of working as the models continue to evolve.

Reach out to our CEO aleksi[at]climatealigned[dot]co to discuss, or drop me a line (founding engineer)

Check some of our past work at https://climatealigned.co or touch base with any of the team with questions on LI

Claude Sonnet 4.6 5 months ago

This whole comment thread here is really echoing and adding to some thoughts ive had lately on the shift from considering LLMs replacing engineering to make software (much of which is about integration, longevity and customization of a general system), vs LLMs replacing buying software.

If most software is just used by me to do a specific task, then being able to make software for me to do that task will become the norm. Following that thought, we are going to see a drastic reduction in SASS solutions, as many people who were buying a flexible-toolbox for one usecase to use occasionally, just get an llm to make them the script/software to do that task as and when they need it, without any concern for things like security, longevity, ease of use by others (for better or for worse).

I guess what im circling around is that if we define engineering as building the complex tools that have to interact with many other systems, persist, be generally useful and understandable to many people, and we consider that many people actually dont need that complexity for their use of the system, the complexity arises from it needing to serve its purpose at huge scale over time. then maybe there will be less need for enginners, but perhaps first and foremost because the problems that engineering is required to solve are much less if much more focused and bespoke solutions to peoples problems are available on demand.

As an engineer i have often felt threatened by LLMs and agents of late, but i find that if i reframe it from Agents replacing me, to Agents causing the type of problems that are even valuable to solve to shift, it feels less threatening for some reason. Ill have to mull more.

I am continually surprised by the reference to "voluntary actions taken by companies" being brought up in discussion of the risks of AI, without some nuance given to why they would do that. The paragraph on surgical action goes in to about 5-10 times more detail on the potential issues with gov't regulation, implying to me that voluntary action is better. Even for someone at anthropic, i would hope that they would discuss it further.

I am genuinely curious to understand the incentives for companies who have the power to mitigate risk to actually do so. Are there good examples in the past of companies taking action that is harmful to their bottom line to mitigate societal risk of harm their products on society? My premise being that their primary motive is profit/growth, and that is revenue or investment dictated for mature and growth companies respectively (collectively "bottom line").

Im only in my mid 30s so dont have as much perspective on past examples of voluntary action of this sort with respect to tech or pre-tech corporates where there was concern of harm. Probably too late to this thread for replies, but ill think about it for the next time this comes up.

My understanding is that modern mobile phone cameras do heaps of "stacking" across multiple axes focus, exposure, time etc to compose a photo that saves onto your phone. I believe its one of the reasons for the multiple cameras on most flagship phones, and then each of them might take many "photos" or runs of data from their sensors per "photo" you take. id love to see a good writeup of the process, but my gut says exactly what they do under the hood would be pretty "trade secret"ie.

For small models this is for sure the way forward, there are some great small datasets out there (check out the tiny stories dataset that limits vocab to a certain age but keeps core reasoning inherent in even simple language https://huggingface.co/datasets/roneneldan/TinyStories https://arxiv.org/abs/2305.07759)

I have less concrete examples but my understanding is that dataset curation is for sure the way many improvements are gained at any model size. Unless you are building a frontier model, you can use a better model to help curate or generate that dataset for sure. TinyStories was generated with GPT-4 for example.

Totally agree, one of the most interesting podcasts i have listened to in a while was a couple of years ago on the Tiny Stories paper and dataset (the author used that dataset) which focuses on stories that only contain simple words and concepts (like bedtime stories for a 3 year old), but which can be used to train smaller models to produce coherent english, both with grammar, diversity, and reasoning.

The podcast itself with one of the authors was fantastic for explaining and discussing the capabilities of LLMs more broadly, using this small controlled research example.

As an aside: i dont know what the dataset is in the biological analogy, maybe the agar plate. A super simple and controlled environment in which to study simple organisms.

For ref: - Podcast ep https://www.cognitiverevolution.ai/the-tiny-model-revolution... - tinystories paper https://arxiv.org/abs/2305.07759

Im not sure about how this translates to react native, AFAICT build chains for apps less optimiside, but using vercel for deployment, neon for db if needed, Ive really been digging the ability for any branch/commit/pr to be deployed to a live site i can preview.

Coming from the python ecosystem, ive found the commit -> deployed code toolchain very easy, which for this kind of vibe coding really reduces friction when you are using it to explore functional features of which you will discard many.

It moves the decision surface on what the right thing to build to _after_ you have built it. which is quite interesting.

I will caveat this by saying this flow only works seamlessly if the feature is simple enough for the llm to oneshot it, but for the right thing its an interesting flow.

I often find that the hard part of writing big, persistant, code is not the writing but the building of a mental model (what the author calls "theory building"). This challenge multiplies when you are working with old code, or code others is working on.

Much of my mental space is spent building and updating the mental model. This changing of my mental model might look like building a better understanding of something i had glossed over in the past, or something that had been changed by someone else. Or it is i think the fundamental first step before actually changing any lines of code, you have to at least have an idea for how you want the mental model to change, and then make the code match that intended change. Same for debugging, finding a mismatch between your mental model and the reality as represented by the code.

And at the end of the day, AI coding tools can help with the actual writing, but the naive vibecoding approach as noted is that they avoid having to have a mental model at all. This is a falacy. They work best when you do have a mental model that is good enough to pass good context to them, and they are most useful when you carefully align their work with your model over time, or use them to explore/understand the code and build your mental model better/faster/with greater clarity.

I think the depth (in time, and community involvement) is one of the things that has drawn me to this project. It has the excellent vibe of a dedicated and yet accessible, IMO because of the beautiful and widely available visual output, internet community.

Thanks for sharing some of this rich history!

cool to hear that the actual electric sheep project is still something you can interact with!

For those super new to it (like me), check out https://electricsheep.org/ og video we came across it with https://www.youtube.com/watch?v=O5RdMvgk8b0 (this was just the first I found when looking there are many on youtube) and the algorithm behind it https://flam3.com/ This is all AFAICT as someone who's only just skimmed the surface, but i find it amazing.

The electric sheep always intrigued me so much! but was a bit before my time, and also felt so impenetrable. I appreciate you drawing the link between them and something like this which is so finite and understandable.

and to OP for making something so finite and understandable ofc.

see my separate comment for more on hackernews.coffee in particular, we (same team, different experiment) are thinking a lot about personal content, and how you have maximum visibility and control.

Keeping these projects separate allows us to test ideas that orbit around a theme (not 100 % sure what the theme is yet, but it features personal, anti-slop content, while still using llms.)