HN user

Jimmc414

14,948 karma

x.com/JimMcM4 github.com/jimmc414

Posts1,270
Comments955
View on HN
www.npr.org 2h ago

ICE shared Medicaid data it wasn't supposed to have with Palantir

Jimmc414
80pts12
spectrum.ieee.org 3h ago

Digital Surveillance Reshapes Fishery Enforcement in Indonesia

Jimmc414
3pts0
arxiv.org 3h ago

SuperPass: Fast-Tracking Blocking Threads to Mitigate Priority Inversion

Jimmc414
3pts0
www.themarshallproject.org 3h ago

State Reverses Course, Finds Cuyahoga Jail Staff Failed to Start CPR

Jimmc414
2pts0
www.npr.org 4h ago

Deloitte systems denied Medicaid to the disabled. New laws could make it worse

Jimmc414
6pts0
www.404media.co 5h ago

Emails Reveal Why a Town Put Bags over Its Flock Cameras

Jimmc414
20pts3
www.gao.gov 6h ago

Critical Minerals: Reducing U.S. import reliance with substitution and recycling

Jimmc414
45pts70
www.propublica.org 6h ago

A Puerto Rico Government Agency Exposed 1M Social Security Numbers

Jimmc414
18pts3
www.404media.co 6h ago

You Opened a Credit Card. ICE Now Knows Where You Live

Jimmc414
20pts21
fasterthanli.me 19h ago

Making Our Own Spectrogram

Jimmc414
4pts0
www.quantamagazine.org 19h ago

Martin Picard's Mitochondrial Theory of Mind

Jimmc414
3pts0
arxiv.org 21h ago

SEAM-V: A Hybrid-Decoupled RISC-V Vector Processor

Jimmc414
3pts0
www.eff.org 21h ago

An Explosion of Surveillance Towers Is Coming to US Borders, Costing $1B+

Jimmc414
13pts0
arxiv.org 21h ago

Hardware Mechanisms to Dynamically Throttle AI Performance

Jimmc414
2pts0
www.themarshallproject.org 1d ago

Amid Increased Scrutiny, ICE Detention and Deportation Data Goes Dark

Jimmc414
144pts37
www.404media.co 1d ago

ICE to Pay Thomson Reuters $125M to Find Voter Fraud

Jimmc414
234pts146
kffhealthnews.org 2d ago

Pregnant Womans Roadside Death Triggers Push to Reopen Mississippi DeltaHospital

Jimmc414
10pts0
papersplease.org 2d ago

Did California's DMV Director lie to the legislature?

Jimmc414
3pts0
www.404media.co 2d ago

New Orleans Cops Published Policy Document Allowing Weaponized Drones

Jimmc414
12pts0
www.statnews.com 2d ago

My husbands suicide shows theres something very wrong with US insurance industry

Jimmc414
153pts163
arxiv.org 2d ago

The Cost and Network Limits of Space-Based AI Compute

Jimmc414
3pts0
arxiv.org 2d ago

Don't Predict, Prioritize: Rethinking GPU Reliability Assessment

Jimmc414
1pts0
www.gao.gov 3d ago

Identity Verification: GSA Needs to Address Fraud Threats and Technical Issues

Jimmc414
3pts0
www.gao.gov 3d ago

Aviation Cybersecurity: Key Shortfalls in FAA/TSA Collaboration on Cybersecurity

Jimmc414
2pts0
arxiv.org 3d ago

Campaign Diagrams: Visualizing the March Through the Phases of a Workload

Jimmc414
1pts0
www.theguardian.com 5d ago

'Adversarial clothing': Garments designed to confuse facial recognition systems

Jimmc414
4pts0
martinfowler.com 5d ago

The Archaeologist's Copilot

Jimmc414
1pts0
papersplease.org 5d ago

If I have a US passport, do I need another ID to enter the US?

Jimmc414
5pts0
spectrum.ieee.org 5d ago

Atomically Thin Materials Significantly Shrink Qubits (2022)

Jimmc414
13pts1
en.wikipedia.org 5d ago

Flowers for Algernon

Jimmc414
9pts3

The government's stated remediation for the Palantir transfer was that the data had been shared over a Microsoft Teams chat and was deleted from the chat. That is not a deletion.

According to CMS's security program documentation, Medicare and Medicaid are a covered entities under HIPAA, subject to 60 day notice to affected individuals, reporting to HHS OCR, and notice to media outlets when a breach affects more than 500 residents of a state.

Not to mention this violated a standing court order.

https://security.cms.gov/learn/cms-breach-response-handbook

Agreed that seems like a potential conflict that should be scrutinized, but it highlights a real problem in health research. Who else would fund a 40,000 participant 15 year study about egg consumption and Alzheimer’s? How many studies are not conducted because no one is interested in funding them?

I think the value right now is to focus less on external orchestration if at all. trust the (current best) model to do it better than anything you bolt on to the harness. focus your energy on providing clearer specs. I think the optimal spec is a disambiguated (through liberal use of the AskUserQuestion tool) 1 intent, 2, input/output contracts 3 constraints and 4 preconditions. focus on that and get out of the models way. I think of it like this, imagine a person who was not as smart as you was trying to tell you how to do a task. would you want more verbosity and step by step instructions or would you want them to just cut to the chase (ie, what are you trying to do, what are the obstacles, I'll let you know if I have questions).

also let the model verify itself. don't give it an objective that is vague, give it clear exit criterias for goals and let it loop until it gets there so much of the orchestration scaffolding seems like massive technical debt

oddly, I do the opposite of a lot of conventional advice when it comes to models. I use no memory, I think there is something similar to context rot when everything is stored. I like creating markdown files as memory that the model can grep if needed. I also havent found a real use for hooks yet, I have tried but they always seem to get in the way. skills on the other hand are very undervalued. they are so much more powerful than many realize. I used to think agents were where the power was. I think its actually skills. agents are really for context preservation. skills are what increase capabilities

I'm not even talking about quantity of items in memory, I mean dilution of intent. I really love a model with a clean slate and only the items it needs. I fear the memory guides the model in areas that might not be what I want with the current prompt

progressive disclosure is a big one. you can make context available but it is only loaded when needed. like lazy loading for prompt engineering. skills are to be used to instruct the model how to do something specific that is not in its training data. like how to access my proprietary system, how to interface with a custom program. you can embed templates in skills, you can embed code that executes in skills and only the output is loaded into context. skills expand capabilities, agents constrain context

(constraining context is a very good thing btw, don't mean to infer that agents are somehow inferior to skills)

I don't see any ethical connection between adding canary tokens to your output to catch people breaking your accepted ToS through ongoing distillation and stealing your PII off of your machine. How are legitimate users in US or China possibly harmed by Anthropic silently changing the apostrophe in Today's or the date separator from - to /?

Respectfully, the story you referenced is about how MSG compiled "Facial Recognition Activists.docx," collecting its critics' tweets and comments into an internally accessible file.

https://www.404media.co/madison-square-garden-made-dossier-o...

This story is about how the hackers actually got into Madison Square Garden.

https://www.404media.co/how-hackers-broke-into-madison-squar...

I think there is an argument that there are two different stories here.

“On one end £9 of labour cost for a plate of asparagus seems deeply inefficient and unrealistic”

Every time I’ve had this thought that I could recreate a dish and spend a lot less, I end up paying more for the same dish, with lower quality and I still have to do the dishes. It might not be as unrealistic as it seems if you aggregate the wages of all staff involved and divide it by the number of plates they serve over a night. I’m not a cook, but if someone offered me $9 to prepare them a comparable asparagus dish, it doesn’t sound that lucrative.

Agreed. It touches on the same points raised in this recent post,

“If You are Asking for Human Attention, Demonstrate Human Effort”

https://news.ycombinator.com/item?id=48497609

If this interview isn’t important enough to assign a human to the task, it sends a message about what you should expect as an employee. I’d expect a larger percentage of these AI interviewers will be vetting AI interviewees.

The reason there’s no good orchestration layer above the base model is that the useful part of orchestration moved into the model. what’s left to bolt on top is negative value when token consumption, latency and usability are factored in.

I’ve spent a great deal of time trying to beat plain Claude Code and Codex with a variety of planner/critic/decomposition orchestration setups. Every one produces less or same quality, 2x+ tokens, more latency and a second system to debug. External orchestration seems like I’m just creating technical debt and the actual value is focusing exclusively on providing clearer specifications (ie statements of intent, constraints and input/output contracts) to the model and just letting the leading frontier model do its thing with reasoning and effort maxed out.

Ironically the real orchestration benefit is on the human side refining intent to model instead of orchestrator to agent, invoking AskUserQuestion and other tools for the model to more deeply disambiguate my requests.

Aikido Code Audit 1 month ago

“But it appears 1 or more organizations have successfully jail-broken Fable 5”

This is hardly true or it’s true of all frontier models and this was only magnified by Fables capabilities. It’s that you could hand Fable 5 vulnerable code, ask it to fix it, return patch plus test cases proving the fix and exploit relevant detail falls out as a byproduct of legitimate secure code review work.

I challenge anyone to provide a fix for this “exploit” without compromising Fable’s ability to patch unsecure code.

Norway spent two decades digitizing classrooms and is now unwinding it. Seems a bit shortsighted and reactionary although I think they are trying to do the right thing.

Plus "Generative AI" isn't one single thing. Using it to write your essay is cognitive offloading but using it as a Socratic tutor that gives immediate feedback and adapts to the student is closer to the thing education research says works.

There's an equity angle as well. A school ban doesn't ban AI at home. It bans the equalizing version. Kids in educated, rich households will get AI exposure from parents. Kids without that won't get it anywhere, because the one place where the field is leveled has opted out. If AI fluency becomes a differentiator in the labor market infrastructure which is very likely a 7 year exposure gap sorted by household class is the opposite of what public education is supposed to be for.

(edit: By AI fluency I mean basically knowing how to drive the tools, an intuition for what the tools can and can't do, when to use AI vs doing it yourself, plus detecting when output is wrong, knowing what to verify, etc.)