HN user

rooftopzen

145 karma
Posts9
Comments44
View on HN

Cliche topic - from a few years ago (the "RAG is dead" vs "All You Need Is Advanced RAG" BS - it came in waves and cycles, spread by bots on social media networks).

"Pruning RAG Context" is trying to recycle the old stuff (again), presuming the reader is naive (implies kapa.ai is not going anywhere). The current cycles were "openclaw" (I think that died), now we are on "harnesses" - when that dies the paid social media bots will give you something else. Shell game.

Just declare / define dictionary as a variable in your prompt to carry forward (when you decide to continue using LLMs for certain things). Also either summarize or truncate history. 3-4 year old concept. Not a big thing.

Great viz; the original paper wasn't peer reviewed; it might be great but I've learned its a waste of time to read those in current times (sorry and take this as one data point that suggests they should have done that).

That said, I've found FAISS great for certain use cases; wanted to say thx for surfacing - its not updated to work with most packages these days outside of faiss-cpu - curious why Meta dropped its maintenance; was it due to its slower speed or otherwise priorities?

Considering how YC companies are customers of other YC companies (presumably to lift ARR), how many YC companies have Delve compliance?

Should we worry about AI startup customer data…

Ezra Klein - any premise he has had in past 1-2 years is built on a maligned belief in AGI that is unscientific. Have heard he recently pivoted eg “moved goal posts” for his AGI timeline but the guy appears to be wasting time following delusional trends, this piece included.

I've spent about an hour a week on this since Jan. Traced a large % of bogus news stories this year back to Reuters (fwiw) before they are picked up by other outlets and spread.

I've found legitimate stories also sourced from Reuters, but haven't found illegitimate stories NOT sourced from Reuters (in other words, they seem to originate from the same source, not sure why)

Abundant Intelligence 10 months ago

Sorry but your concept of AI is marketing driven. It's probabilistic, understanding is past your pay grade.

Abundant Intelligence 10 months ago

Exactly - AI allows for intersections in concepts from training data; up to the user to make sense of it. Thanks for stating this (I end up repeating same thing in every conversation, but is common sense).

Abundant Intelligence 10 months ago

Sam Altman has been a joke for awhile now, heard only his investors defend him for their next round increase - is that who you are?

You naively replaced deterministic process w probabilistic process - following a trend that is uneducated.

I am taking screenshots of blogposts like this for a museum exhibit opening next year - lmk if you’re willing.

Caveats: 1. It's 101 pages (do # of pages correlate with aggressive effort to be authoritative, e.g. 'state of the world' pdfs. 2. This appears to come from the AGI existential threat doomer camp - does it even have any validity? At first glance, it appears both absurd in terms of presumptive risks, and also AI generated (this is biased to the pages I'd read) 3. MLCommons has a more scholastic approach to ranking models on potential harms, curious on reception to all approaches (pros vs cons of each)

Lol

Yh I get your point - post is not necessarily designed to prove AI use (it's already highly probable, and not necessarily bad by itself in theory) it's the implications of it that are more interesting than deterministic evidence of it, but by showing evidence of it being likely - updated the post to reflect a better baseline.

Not following exactly, so apologies if I'm misinterpreting, but I'm the author and updated this post (transparently) with nuance I'd recently learned about that explains this (somewhat) - the larger bills contain entire pages with only headings that contain emdashes - removed the headings from analysis so that the emdashes per page are only from the legislative text itself. For the conservatively / minimal difference, we're still looking at a 30% increase from a decent baseline.

I'm the author and updated this post - after looking into this, the larger bills contain entire pages with only headings that contain emdashes - removed the headings from analysis so that the emdashes per page are only from the legislative text itself. For the baseline, over 50% of bills found on congress.gov are 1-2 pgs, after reading a few I decided some rationale could exist to remove them from the baseline - even after all these adjustments, we're still looking at a 30% increase from a decent baseline of similar bill size. It's evident when reading the text below headings (as a human!)

If I’m following correctly, this drama is like a Netflix series (and I agree, crazy stuff we couldn’t make up). If it’s only the Trump admin policies here, yeah it’s crazy bold (and which I state the ethical implications of).

Science fiction plot twist could involve anything completely crazy in the bill no one notices by having to use an LLM to read and that is open to interpretation enough to be only decided by the courts later.. I didn’t look for anything hidden and vague; but how would one really know lol.

Share IT is from 2024, but the 2017 tax cut bill is interesting (lots of emdashes there that deviate from the avg) - you’re correct on the additional need for text analysis in this case. Bills I’d found from earlier in 2024 that are publicly available do not have emdashes outside of the table of contents, which is built into the average - curious how/why they are used so much in this bill from 2017, now wondering how they got into any potential templates (or not), and adds the confound of how much this is AI or template (or requirements, or something else) Thx!

First, see below for Toll et al 2020 and I used autocorrect for grammar. Sorry you were dismissive before looking it up, is more a reflection of your bias.

https://liu.diva-portal.org/smash/get/diva2:1591409/FULLTEXT...

Second, I noted all caveats with an LLM counting that - I actually presumed I undercounted, but it had been noted that a simple ctrl-f found 3.8 per page rather than 9.8 per page (counting only single emdashes not double). The actual number doesn’t matter so much, since low bound is absurd difference from baseline bills I checked from earlier this year and 2024, where they do not exist outside of the table of contents.

4.x emdashes per page (low bound) is absurd, and the implication of this is the point you (respectfully) missed.