Sure, but it doesn't really fit there as a joke, it looks like it's just meant to be part of what they were trying to say.
HN user
LiamPowell
AI Effect
Honestly? That's not just valuable—it's essential.
I'm curious if you wrote this or had a LLM write it.
I'm genuinely curious to be clear as I don't see why anyone would bother to go through a LLM to write such a short reply. Have we reached the point where Claudeisms that are this obnoxious have become part of regular speech?
I suspected as much, and that brings us to the second issue where if we use a cohort of judges then the model that likes it's own code the most still wins.
Here's the question I ask about every project that claims to make a LLMs output so much better: If it works so well then why would the model provider not just put it in the system prompt? Or in the case of interactive skills, why would Claude Code/Codex not make it a core part of the product?
On top of that, if your magic markdown file really does work then where's the evidence showing that? These projects never include even basic benchmarks. At best they're entirely vibe based, however more often they're completely untested. Give us a proper benchmark, even a single prompt and it's output with and without your skill in use would be better than every other project out there.
This is not actually what the reviewer prompt says, or perhaps it is, I don't know since they don't make it public. I'm just pointing out how it seems like a bad idea to ask a LLM to make a subjective judgement on things like "taste". If the SOTA LLM witting the code could not produce tasteful code then why would a different LLM be able to judge the "taste" of that code?
Which LLM should we even use to judge taste? Is it giving an unfair advantage to Model X if we use Model X as the judge? Maybe we should use multiple models as the judge, but now the model that's best at recognising and praising its own code has an advantage. The whole thing is just an unsolvable problem when a LLM is the judge.
You are a senior SWE-Bench reviewer, make no mistakes.
I don't know what a better approach would look like while still remaining feasible, however this approach of telling a LLM to make a subjective judgement seems fundamentally flawed.
I'm not sure about Kalshi, however on most sports betting sites you actually are betting against the house. The betting sites all have in-house models (or piggyback off other sites) that are much better at predicting odds than the general public. If someone is making money then the sites just place limits on that account so they're not losing money.
Most ad blockers do already use MV3, uBlock Origin is the only one still using V2 as far as I know.
There are some drawbacks to V3, however none prevent creating an effective ad blocker, as demonstrated by the fact that many exist. Though saying that doesn't make for nearly as effective clickbait...
OP, I assume your comment[1] is getting flagged because of the obvious LLM usage. No one wants to interact with a comment that's not written by a human.
That don't fall back to Opus if their classifier thinks you might be working on anything that might be a competitor's product. It silently injects instructions into the prompt to sabotage your work. Read the policy above, it's insane to me that they're publicly admitting to this.
The assumptions are so much worse than that:
Methodology & assumptions: No caching
This is absolutely absurd. Claude code is of course using the cache (and this can be verified by looking at the traffic). It would be an incredibly stupid design to resend the whole input without a cache for every input, every tool use, etc..
especially with all the stuff that SpaceX has put into orbit in recent years
I've heard this repeated a lot but I've never seen anyone do the maths. StarLink satellites are all in very low orbits, so intuitively it seems like most debris from a collision would just end up deorbiting.
Maybe, but they certainly used it for marketing too. At the time they contacted a bunch of publications and gave them access but told them they could only share snippets of the output [1]. The only reason to set restrictions like that is marketing.
[1] https://youtu.be/TfVYxnhuEdU?t=102
Transcript of the timestamped part:
Now, OpenAI's terms of service don't let me give you the full list. I have to curate them, and show you a sample. Those are the terms and conditions I agreed to.
They did it for 2 and 3, however it looks like they didn't for 4 and 5.
GPT-2: https://slate.com/technology/2019/02/openai-gpt2-text-genera...
GPT-3: https://www.itpro.com/technology/artificial-intelligence-ai/...
OpenAI has been pulling this marketing trick for years. Remember how GPT-3 was too dangerous to release? It's also probably bad PR if script kiddies have access to GPT model with no guardrails even if it doesn't enable any significant attacks.
I don't think I've ever seen LLM output as bad as this output. They sometimes write like that, but not every second sentence.
What's this nonsensical video on the product page that allegedly shows an "all new thermal system"? https://videos.ctfassets.net/jy9s7k22hbg4/44R1LH71xb8uO4c9dD...
LLMs are not yet capable of generating the level of marketing wankery seen here.
TLDR:
SQLite does not (currently) accept agentic code. However the project will accept agentic bug reports that include a reproducible test case. Patches or pull requests demonstrating a possible fix, for documentation purposes, are welcomed.
Extensions never had to be given unsandboxed access to everything. That's a choice that they actively made.
When did KVM switches get so expensive? Level1Techs doesn't appear to be much more expensive than the competition, but the margin on all of these has to be absurd. They're not a particularly niche product and the BOM cost is only going to be $20 at most (a TMUXHS4612 is $1 for reference).
I'm amazed that there's not more competition bringing the price down here.
Trivially the answer is yes by the infinite monkey theorem. If we allow the sampler to pick any token then any stream of arbitrary tokens can be generated. Therefore if an original idea can be represented with written words then a LLM can generate it. That is perhaps not the most satisfying answer, but if you want a better one you'll need to provide a function that determines if an idea is original.
Why are "premium" laptop vendors still putting vents on the bottom of their machines? Did they never try actually putting their laptop on their laps and realise how much that design sucks?
2×? Try 5× for the Noctua NF-A12x25 compared the the Arctic P12 Pro that matches or beats it in most metrics. Which isn't to say the Noctua fan is bad, it's just a luxury product for reasons other than performance.
Last I checked they weren't really any quieter than their competitors at the same airflow and pressure (which is a little subjective because your curve will never match perfectly). They do have a really low number on their specs because they have a really low max RPM, but that's not really relevant when you can just lower the speed of other fans.
They're still really good fans, but a lot of this is just marketing.
At max power the Noctua NF-A12x25 has 56 CFM and 2.3 mmAq for 31dBA [1]. At 70% the Artic A12 Pro is 56 CFM, 4.3 mmAq, and 31dBA [2]. At 60% the Asus ProArt PF120 is 61 CFM, 2.6 mmAq, and 30 dBA [3].
Note that the ProArt is a bit thicker (25 vs 30 mm) and all these dBA numbers are almost certainly unobstructed airflow. The Noctua is certainly good, but it's literally over 5× the price of the Artic.
[1] https://www.cybenetics.com/evaluations/fans/4/
It's par for the course in the premium PC parts industry. It's overkill in a way that does not impact performance at all because gamers will pay for that.
The very simplified answer is that the models are first trained on everything and then are later trained more heavily on golden samples with perfect grammar, spelling, etc..
This has come up multiple times before [1], and more generally it's come up hundreds of times with Unix style tools in general. It's always been a stupid idea for every tool to have its own barely documented file format.
This wouldn't be an issue if patches were XML or JSON with a well defined schema, but everything must be a boutique undocumented format in the world of Unix tools.
Maybe the worst part about this is that it can entirely come from a patch being exported by git and then imported straight back in to git. If you can't even handle your own undocumented format then what hope do other tools have that want to work with it?
I can not figure out what on Earth they've done with these graphs, it almost seems like these are an artists impression of a graph.
Looking at the commit graph: Why do commits have big steps followed by slow rolloffs? Why do the steps not happen at uniform points Why do larger steps sometimes have less of a slope than smaller steps but not all the time?
Then looking at the other graphs there's completely different effects going on.
Those are legitimate grievances as mentioned, what they are not is Palantir themselves collecting massive amounts of data, which is often what they're portrayed as doing and what the GP asked about.