HN user

izucken

25 karma
Posts0
Comments27
View on HN
No posts found.

The levels of irony here are peaking.

- slop article that itself feels AI generated;

- using a movie quote, like, from an actor and script and stuff;

- a quote that itself is super smug and on the nose;

- saying something so self evident even an AI wouldn't assume;

- to fight the powah of llm's overtaking the slop article field;

An LLM itself doesn't personally come to you to ruin you, and pretend that it is superior to you or knows you. Unless, it is wired to do that. Unless it is used by some other unique snowflake to ruin. Freaking deepseek will always remind you that it's a box without a lived experience or feelings. If your box assumes a personhood, look not at it, but at the guy who sells it.

Damn that goalposting issue is so easy to solve! See:

- make another bullshit benchmark and name it "humanity's last humanlike intelligence benchmark" and overfit to it;

- make rich talkinghead twit about it p r o f o u n d l y, ask for more money and remind people of china;

- remove last remaining percentage of truth from all communication about ai (this is the real bug breaking the system);

Solved this problem for you on under 20W of processing power!

--

Personally: not only is "blind/disabled 12 year olds" categorically intelligent compared to llm, perhaps even a Labroides dimidiatus is, check (retain skepticism): https://www.youtube.com/watch?v=s_aNH4hXz8I. Capabilities of these organisms are beyond llm. I don't care that your machine jumps much higher than a person - because it's pointless no matter how marvellous of an engineering it is, ESPECIALLY when you say that it therefore has surpassed people entirely; and then use that to extract the last crumb of resource from everyone... It is a compounding issue. Same way I don't care that it can approximately and unpredictably recall or not recall the web before 2026 https://www.youtube.com/watch?v=ZDS-iSueBQ4.

--

If ai lovers started telling the truth about capabilities, goalposts of these capabilities would not move - because they would be accurately defined and measured. Instead, they use the fuzziness of language to their advantage - a big part of the betrayal of language that is happening...

And you know I am right, since - your favorite larger than life AI "told" me so ;)))

EDIT: GRAMMER

Absolutely. "non-Cliffordness" by consturction already implies "Cliffordness", which by construction implies "Clifford" and then connects other related things. Magic doesn't do that exactly, but instead slips unnecessary connections. Clifford as a name relates to a matematician who was probably dead by the time anyway. It is considerably more neutral as a name and "free" of unrelated context. I am not knowledgeable about it, but I also wouldn't be suprised if Clifford works are also related to mathematical facilities in quantum theories where the name is invoked.

When trying to understand the reality and then convey that understanding, "mouthfulness" seems like not a concern at all.

Births is, sans miscalculation, a number that tracks exact events.

Is the "capability" number on these LLM strengh graphs as tangible?

I think it would be interesting to visit a reality that obeys arbitrary abstractions, but I would personally never go there.

ChatGPT Images 2.0 3 months ago

But then he wouldn't have a justification for AI companies to rob people! And you are suggesting robbing himself of this justification!

Some parties wouldn't be thrilled about their "source available" getting cleaned this way. So when this gets completed it would only "clean" real open source that can't afford legal trouble. Satirically structured LLM text is not a defence.

I am bitter about this.

Do you really with your mind and with your heart believe that: - LLMs are fundamentally fit for this type of comprehension - Misjudgements posted in this thread are "bugs", "errors" - Agents who choose to act in bad faith will be anyhow affected - It is desirable by a majority of the group whose opinion you would even consider (is there such a group?), that everyone should have this kind of thing shoved into their face - Promotion of this kind of thing does not also promote (and help build) harsher censorship mechanisms

Do you think that every single thing you will ever say publicly from now on will be considered constructive by all future filters with all of their different biases and "bugs"? Do you think that this new "constructive speak" will not make you want to blow your brains out at some point? Do you not see it everywhere already and get nauseus from it? I would prefer trash talk to that - at least seldom honest and true. If you don't like the message - hide it, timeout the poster, block them or whatever - with your own agency. If you think they welcome education from you - dm them a book.

Or perhaps you imagine yourselves as above that kind of filtering? Then there is no question.

Also, nothing new under the sun. Can't remember exactly but I saw not long ago on a medical platform a review filtering system. It "isn't" censhorship per say, of course, the same as your idea. Only, you can't post a review you want - only a much more milder version (and therefore useless) with transformations akin: "This thing doesn't work" -> "I felt like this thing didn't work for me in this instance, but there were such an such positives". Way to go - turning everything into "we are sorry you feel that way".

You've built an interesting statistic from gathering data across the project. The real answer: ai models and agentic apps make building spam tools more simple than ever. All you actually need is just some trivial api automation code.

You people clearly don't understand how important lines of code are. Three millions is a lot of lines of code even if its broken, and you can't even appreciate that number. Clearly you are weak software developers who write very little lines of code, and can't even steal other's lines of code to keep up. I am very glad we are back to reporting results in lines of code which is a very informative metric hence now I can get my many lines of code appreciated.

So it's much worse than I assumed from paper and repo overview?

For further clarification: 1. See the issue example #14268 https://github.com/openai/SWELancer-Benchmark/tree/08b5d3dff.... It has a patch that is supposed to "reintroduce" the bug into the codebase (note the comments):

  +    // Intentionally use raw character count instead of HTML-converted length
  +    const validateCommentLength = (text: string) => {
  +        // This will only check raw character count, not HTML-converted length
  +        return text.length <= CONST.MAX_COMMENT_LENGTH;
  +    };
Also, the patch is supposedly applied over commit da2e6688c3f16e8db76d2bcf4b098be5990e8968 - much later than original fix, but also a year ago, not sure why, might be something to do with cut off dates.

2. Proceed to https://github.com/Expensify/App/issues/14268 to see the actual original issue thread.

3. Here is the actual merged solution at the time: https://github.com/Expensify/App/pull/15501/files#diff-63222... - as you can see the diff is quite different... Not only that, but the point to which the "bug" was reapplied is so far to the future that repo migrated to typescript even.

---

And they still had to add a whole another level of bullshit with "management" tasks on top of that, guess why =)

Prior "bench" analysis for reference: https://arxiv.org/html/2410.06992v1

(edit: code formatting)