HN user

LiamPowell

1,382 karma
Posts24
Comments272
View on HN
github.com 2mo ago

Add a Prototype Agents.md File

LiamPowell
5pts2
godbolt.org 5mo ago

CCC (Claude's C Compiler) on Compiler Explorer

LiamPowell
17pts9
blog.liampwll.com 6mo ago

Better Communication Protocols with Ada's Record Representation Clauses

LiamPowell
4pts0
muen.codelabs.ch 7mo ago

Muen – An x86/64 Separation Kernel for High Assurance

LiamPowell
2pts1
www.trendmicro.com 7mo ago

EvilAI Operators Use AI-Generated Code and Fake Apps for Far-Reaching Attacks

LiamPowell
3pts1
www.trendmicro.com 7mo ago

EvilAI Operators Use AI-Generated Code and Fake Apps for Far-Reaching Attacks

LiamPowell
2pts1
gcc.gnu.org 9mo ago

Mutably Tagged Types with Size'Class Aspect

LiamPowell
2pts1
bugzilla.mozilla.org 9mo ago

Password manager should support OS X Keychain (2001)

LiamPowell
2pts2
sparforte.com 9mo ago

The SparForte Programming Language

LiamPowell
4pts1
news.ycombinator.com 10mo ago

Tell HN: EasyList break more than just YouTube

LiamPowell
3pts0
en.wikipedia.org 10mo ago

AI Effect

LiamPowell
12pts1
docs.mimer.com 11mo ago

Module SQL

LiamPowell
2pts1
news.ycombinator.com 1y ago

Tell HN: EasyList breaks a decent chunk of the internet

LiamPowell
2pts1
gcc.gnu.org 1y ago

Mutably Tagged Types with Size'Class Aspect

LiamPowell
1pts1
news.ycombinator.com 1y ago

Could you convince a LLM to launch a nuclear strike?

LiamPowell
5pts5
news.ycombinator.com 1y ago

Ask HN: Has anyone had even minor success with LLM powered code completion?

LiamPowell
1pts4
learn.adacore.com 1y ago

Ada for the C++ or Java Developer – Concurrency

LiamPowell
10pts4
600f3559.prunt-docs.pages.dev 1y ago

Show HN: A 5th order motion planner with PH spline blending, written in Ada

LiamPowell
118pts32
forums.lanik.us 1y ago

SuperMicro IPMI - EasyList Forum

LiamPowell
2pts1
blog.adacore.com 2y ago

AdaCore Enhances GCC Security with Innovative Features

LiamPowell
5pts2
taste.tools 2y ago

Taste: A tool-chain targeting heterogeneous embedded systems

LiamPowell
2pts0
forums.lanik.us 2y ago

_300x250_

LiamPowell
3pts1
github.com 2y ago

Adamant – An Embedded Software Framework

LiamPowell
1pts0
dl.acm.org 3y ago

A question-answering system for high school algebra word problems (1964)

LiamPowell
7pts2

Here's the question I ask about every project that claims to make a LLMs output so much better: If it works so well then why would the model provider not just put it in the system prompt? Or in the case of interactive skills, why would Claude Code/Codex not make it a core part of the product?

On top of that, if your magic markdown file really does work then where's the evidence showing that? These projects never include even basic benchmarks. At best they're entirely vibe based, however more often they're completely untested. Give us a proper benchmark, even a single prompt and it's output with and without your skill in use would be better than every other project out there.

This is not actually what the reviewer prompt says, or perhaps it is, I don't know since they don't make it public. I'm just pointing out how it seems like a bad idea to ask a LLM to make a subjective judgement on things like "taste". If the SOTA LLM witting the code could not produce tasteful code then why would a different LLM be able to judge the "taste" of that code?

Which LLM should we even use to judge taste? Is it giving an unfair advantage to Model X if we use Model X as the judge? Maybe we should use multiple models as the judge, but now the model that's best at recognising and praising its own code has an advantage. The whole thing is just an unsolvable problem when a LLM is the judge.

I'm not sure about Kalshi, however on most sports betting sites you actually are betting against the house. The betting sites all have in-house models (or piggyback off other sites) that are much better at predicting odds than the general public. If someone is making money then the sites just place limits on that account so they're not losing money.

Most ad blockers do already use MV3, uBlock Origin is the only one still using V2 as far as I know.

There are some drawbacks to V3, however none prevent creating an effective ad blocker, as demonstrated by the fact that many exist. Though saying that doesn't make for nearly as effective clickbait...

Claude Fable 5 1 month ago

That don't fall back to Opus if their classifier thinks you might be working on anything that might be a competitor's product. It silently injects instructions into the prompt to sabotage your work. Read the policy above, it's insane to me that they're publicly admitting to this.

The assumptions are so much worse than that:

Methodology & assumptions: No caching

This is absolutely absurd. Claude code is of course using the cache (and this can be verified by looking at the traffic). It would be an incredibly stupid design to resend the whole input without a cache for every input, every tool use, etc..

especially with all the stuff that SpaceX has put into orbit in recent years

I've heard this repeated a lot but I've never seen anyone do the maths. StarLink satellites are all in very low orbits, so intuitively it seems like most debris from a collision would just end up deorbiting.

Maybe, but they certainly used it for marketing too. At the time they contacted a bunch of publications and gave them access but told them they could only share snippets of the output [1]. The only reason to set restrictions like that is marketing.

[1] https://youtu.be/TfVYxnhuEdU?t=102

Transcript of the timestamped part:

Now, OpenAI's terms of service don't let me give you the full list. I have to curate them, and show you a sample. Those are the terms and conditions I agreed to.

OpenAI has been pulling this marketing trick for years. Remember how GPT-3 was too dangerous to release? It's also probably bad PR if script kiddies have access to GPT model with no guardrails even if it doesn't enable any significant attacks.

TLDR:

SQLite does not (currently) accept agentic code. However the project will accept agentic bug reports that include a reproducible test case. Patches or pull requests demonstrating a possible fix, for documentation purposes, are welcomed.

When did KVM switches get so expensive? Level1Techs doesn't appear to be much more expensive than the competition, but the margin on all of these has to be absurd. They're not a particularly niche product and the BOM cost is only going to be $20 at most (a TMUXHS4612 is $1 for reference).

I'm amazed that there's not more competition bringing the price down here.

Trivially the answer is yes by the infinite monkey theorem. If we allow the sampler to pick any token then any stream of arbitrary tokens can be generated. Therefore if an original idea can be represented with written words then a LLM can generate it. That is perhaps not the most satisfying answer, but if you want a better one you'll need to provide a function that determines if an idea is original.

StarFighter 16-Inch 3 months ago

Why are "premium" laptop vendors still putting vents on the bottom of their machines? Did they never try actually putting their laptop on their laps and realise how much that design sucks?

Last I checked they weren't really any quieter than their competitors at the same airflow and pressure (which is a little subjective because your curve will never match perfectly). They do have a really low number on their specs because they have a really low max RPM, but that's not really relevant when you can just lower the speed of other fans.

They're still really good fans, but a lot of this is just marketing.

At max power the Noctua NF-A12x25 has 56 CFM and 2.3 mmAq for 31dBA [1]. At 70% the Artic A12 Pro is 56 CFM, 4.3 mmAq, and 31dBA [2]. At 60% the Asus ProArt PF120 is 61 CFM, 2.6 mmAq, and 30 dBA [3].

Note that the ProArt is a bit thicker (25 vs 30 mm) and all these dBA numbers are almost certainly unobstructed airflow. The Noctua is certainly good, but it's literally over 5× the price of the Artic.

[1] https://www.cybenetics.com/evaluations/fans/4/

[2] https://www.cybenetics.com/evaluations/fans/175/

[3] https://www.cybenetics.com/evaluations/fans/229/

The very simplified answer is that the models are first trained on everything and then are later trained more heavily on golden samples with perfect grammar, spelling, etc..

This has come up multiple times before [1], and more generally it's come up hundreds of times with Unix style tools in general. It's always been a stupid idea for every tool to have its own barely documented file format.

This wouldn't be an issue if patches were XML or JSON with a well defined schema, but everything must be a boutique undocumented format in the world of Unix tools.

Maybe the worst part about this is that it can entirely come from a patch being exported by git and then imported straight back in to git. If you can't even handle your own undocumented format then what hope do other tools have that want to work with it?

[1] https://mas.to/@zekjur/116022397626943871

I can not figure out what on Earth they've done with these graphs, it almost seems like these are an artists impression of a graph.

Looking at the commit graph: Why do commits have big steps followed by slow rolloffs? Why do the steps not happen at uniform points Why do larger steps sometimes have less of a slope than smaller steps but not all the time?

Then looking at the other graphs there's completely different effects going on.