HN user

trunch

113 karma
Posts0
Comments13
View on HN
No posts found.

If you want another story to run, I'd really love to see an investigation into how these different companies are convincing governements that the only path forward to win global dominance is through achieving 'agi' first and how much that contributes to the reckless acceleration of ai software and infrastructure development

Also a good expose on accelerationists and e/accs and who among the elites fall in this group is direly needed as well

Which of the LiveCodeBench Pro and SWE-Bench Verified benchmarks comes closer to everyday coding assistant tasks?

Because it seems to lead by a decent margin on the former and trails behind on the latter

I'm usually very supportive of EU tech regulation, but to be honest I don't really want to put my name and address up on apps I throw up on the store

Would like to keep my identity separate to whatever projects I have usually, especially if they're ones that don't 100% align with the your own developer brand that employers might screen for

I wanted to write about the abandoned kasbah of Foum Zguid yesterday,

I know this sounds trite but please do actually post this I'd love to read it. With this kind of post, if anyone learns about it it's far more likely they've learned about it from your blog than having requested info about it directly from AI.

I've a million misgivings about the future with AI and the people who control and benefit from it, but I do still think one valuable thing that will be able to survive is humans ability to direct attention and bring value to otherwise neglected stories like what you're wanting to write about there.

50+ hours on 256 H100s is considered impressively low training?

Really makes me wonder if any of this incredibly computationally expensive research is worth it, which seems only useful in potentially promising a future in which humans are given less opportunity to express themselves creatively - while delivering them an infinitely produceable amount of ai generated 'content' to passively consume

Because this will definitely be used only to innocently tell off people doing 1/10 the work of everyone else, and not micromanage and hound people to increasingly unrealistic standards in already desperate conditions.

Safe to say you aren't in any position where every move you make will be watched by AI and analysed for faults so that your boss can scream at you more efficiently whenever you don't meet standards for their pitiful wages.

Not OP but other than what core functionality they can demo to investors, every AI company seems to have extremely lacking:

- web design (basic features take years to implement, and when done break the website on mobile)

- UI/UX patterns (cookie cutter component library elements forced into every interface without any tailoring to suit how the product is actually used, also makes a Series C venture indistinguishable from something setup in a weekend)

- backend design (turns out they've been hemorrhaging money on serverless Vercel function calling instead of using Lambda and spending a minute implementing caching for repeat requests)

- developer docs (even when crucial to business model, often seems AI generated, incomplete, incoherent)

And this usually comes from hiring much less developers than is needed, and those that are hired are 10x Cursor/GPT developers which trust it to have done a comprehensive job at what seems like a functional interface on the surface, and have little frame of reference or training for what constitutes good design in any of these aspects.

Intrigued by the project, love the idea of exhaustively exploring significance for arbitrary input.

Think you need to provide a few example inputs and outputs on the github for the program.

Also not sure a project focused on decoding meaning and signals benefits from having AI generated interpretations divorced from the inherently human act of sign interpretation. Can be seen in the md file, such a rigidly enforced structured output has forced it to give some averaged amount of weight to different categories and examples, when many are facets of each-other, or purely just expressions of something mentioned earlier. I can see free-form high-temperature llm outputs fed to another model, which serves only to aggregate their core interpretations, providing more insight than what's within the document currently.