HN user

doctoboggan

6,655 karma

Owner, Lulim Jewelry (https://lulimjewelry.com)

http://jack.minardi.org

contact me at:

    python -c "print '{}'.join(['jack', 'minardi', 'org']).format('@', '.')"
[ my public key: https://keybase.io/jminardi; my proof: https://keybase.io/jminardi/sigs/UeFUbvxi2yZJWRicpP5aAtgs0AK0mzQytx2Zefu2RM8 ]
Posts62
Comments1,046
View on HN
www.axios.com 6d ago

Teleprompter operator investigated over alleged Kalshi trades on Trump speeches

doctoboggan
5pts0
www.youtube.com 9d ago

Discovering so many hidden sounds With Fotric Acoustic Imager [video]

doctoboggan
3pts0
getnlab.com 1mo ago

NLAB: The worlds smallest electronics lab

doctoboggan
14pts13
www.technologyreview.com 3mo ago

"Constellations: A story about seeking" by Jeff VanderMeer

doctoboggan
1pts0
www.youtube.com 4mo ago

Analysing the First Solid-State Battery [video]

doctoboggan
2pts1
9to5mac.com 4mo ago

Claude hits #1 on the App Store as users rally behind Anthropic

doctoboggan
105pts2
9to5mac.com 5mo ago

Apple reportedly pushing back Gemini-powered Siri features beyond iOS 26.4

doctoboggan
4pts0
riviantrackr.com 7mo ago

Rivian Unveils Custom Silicon, R2 Lidar Roadmap, and Universal Hands Free

doctoboggan
395pts647
www.youtube.com 11mo ago

Is Information a Fundamental Force of the Universe? [video]

doctoboggan
3pts0
techcrunch.com 1y ago

Elon Musk's SpaceX might invest $2B in Musk's xAI

doctoboggan
5pts1
www.youtube.com 1y ago

First Autonomous Delivery of a Car – Tesla [video]

doctoboggan
10pts3
www.youtube.com 1y ago

Expanding Racks [video]

doctoboggan
135pts17
gizmodo.com 1y ago

Hugo Administrators Resign in Wake of ChatGPT Controversy

doctoboggan
13pts2
www.youtube.com 1y ago

One Step Closer to a 'Grand Unified Theory of Math': Geometric Langlands [video]

doctoboggan
2pts0
iac.gatech.edu 1y ago

Can AI make ethical decisions?

doctoboggan
2pts0
www.youtube.com 1y ago

Watch Live: Kilauea Volcano Erupts on Hawaii's Big Island [video]

doctoboggan
15pts5
www.youtube.com 1y ago

2024's Biggest Breakthroughs in Physics [video]

doctoboggan
1pts1
news.ycombinator.com 1y ago

Ask HN: Anyone using an LLM to filter email?

doctoboggan
1pts3
www.youtube.com 1y ago

What Your Brain Is Doing When Doing 'Nothing' [video]

doctoboggan
2pts1
www.youtube.com 2y ago

The solar-powered aircraft flying high in the atmosphere [video]

doctoboggan
1pts0
www.youtube.com 2y ago

Scott Manley: SpaceX Orbit Largest Spacecraft in History [video]

doctoboggan
27pts3
status.quay.io 2y ago

Red Hat Quay.io "Sporadic image pull failures"

doctoboggan
1pts1
news.bloomberglaw.com 2y ago

Biden Targets Artificial Intelligence in Broad Regulation Order

doctoboggan
4pts1
www.youtube.com 2y ago

Are Room Temperature Superconductors impossible? [video]

doctoboggan
2pts0
www.wsj.com 3y ago

Elon Musk Is Making Mark Zuckerberg Seem Cool Again

doctoboggan
5pts9
old.reddit.com 3y ago

iOS Reddit App Apollo's Developer Surprised by WWDC Callout

doctoboggan
133pts27
www.yahoo.com 3y ago

MSCHF’s 'Tax Heaven 3000' Is a Dating Simulator That Also Files Your Taxes

doctoboggan
172pts42
www.youtube.com 3y ago

GPT Developer Livestream [YouTube]

doctoboggan
2pts0
arxiv.org 3y ago

Study: Do Users Write More Insecure Code with AI Assistants?

doctoboggan
4pts2
news.ycombinator.com 3y ago

Ask HN: Tips for using OpenAI’s lower end models

doctoboggan
9pts3

Similar to the story of George Dantzig, who was late to class and solved two open problems in statistics because he mistook them for homework, I think the current batch of frontier LLMs are chained up by knowing which problems are supposed to be unsolved. If they're let free (probably via some targeted RLHF) we might get a flurry of solutions to open problems.

The legal argument is that we gave this data "voluntarily" to the data broker so the government no longer needs a warrant.

This is an area that desperately needs new laws to catch up to the reality of what is going on, but I don't really see much motion toward that goal in the near future.

I have a side business selling custom fingerprint jewelry and I use gemini nano banana to clean up customer submitted fingerprint images. This was a step I used to do by hand at 10 - 15 minutes per image and nano banana is the first model that is able to do the task (it is astonishingly good at it). I can't wait to see what the next nano banana can do, hopefully its released soon.

I think the SotA is moving too fast for the production timelines of an ASIC, wouldn't you think? People are just now coming out with LLAMA ASICS but who would want to use LLAMA? Or I guess you are arguing that the models _now_ will be durably useful enough to commit the time to creating the ASIC?

What are the steep angle arrows indicating? Too steep for an escalator/stairs, but not 90 degrees like an elevator would need to be. Anyone know?

Is there anyone credible who thinks this is a plausible pathway for SpaceX to make huge amounts of profit?

Scott Manly (who I think is credible) has a video where he goes over the logistics of SpaceX's space based data centers. He seems to think its an idea worth pursuing, but its important to note that his expertise is space tech, and not business strategy.

Can you tell me more about the cache misses causing a hefty bill? I think I read somewhere that interacting with a CC instance that has been idle for over an hour can cause a cache miss. Is this what you are referring to? How hefty of a bill are we talking? (using Fable for instance)

I do not mind when I am coding with Claude and it uses all the typical claudisms. I am much more bothered when I am reading a blog post, email, or other form of prose and I see those same claudisms.

I guess they are not annoying since I know I am talking to an LLM and expect the typical responses. When I am reading prose online that I previously would have expected a human to write, it can be quite jarring to realize its an LLM.

Very interesting, I wonder what happened in 2020 that causes the rotational speed to start drifting the other way?

Pandemic -> more people working from home -> less people in tall office buildings -> faster rotation (like a skater pulling in their arms).

Probably not remotely true but it would be funny.

You can never ask why a model did a certain thing, or what it was "thinking" when it said something - just like you can't ask a human which neurons were firing when they had a certain thought. The information just isn't available at that level.

You absolutely can have deep nuanced discussions with LLMs however, you just need to better understand their strengths and weaknesses.

Yes, good science writing almost always gets an opinion from someone not involved in the research for the article. I would guess varying definitions of "not involved" depending on the repute of the publication.

I don't think you can get to $68 for half as many people, even with drinks and tax.

A 5pc chicken tenders, Mac and cheese, and a large drink is $25 before tax. If there are three people who get a similar meal (but not exact so they don't share the family meals) then the total is $75 before tax. Seems like the original price quote of $68 is certainly plausible for a group of three. I am sure its possible to feed three people for less like you claim, but that doesn't mean the $68 is impossible to reach.

Interesting project and (possibly more) interesting explanation of the development process. I agree with the author that the primary difference between vibe slop and real engineering is just reading the lines of code. However it does feel like we are just on the cusp of only needing to read the tests and _not_ all the lines of code. Maybe a few more model generations and we will be there.

Claude Sonnet 5 22 days ago

No, you are misunderstanding the graph. Draw a vertical line anywhere, that is a "constant cost" line. For any given cost, Opus 4.8 has a higher performance than Sonnet 5. Only where Sonnet 5 effort is at medium or low would it make any sense to use it, as there isn't even an equivalent Opus effort level to compare to.

Alternatively you can draw a horizontal "constant performance" line and see that Opus is cheaper for a given performance level.

Claude Sonnet 5 22 days ago

The cost per task chart is telling me that I should _never_ use Sonnet 5 above medium effort level - Opus always performs better for a given cost. So I guess the takeaway is that if Sonnet 5 medium isn't good enough for you, switch models, not effort levels.

If the Chinese government is as involved in LLM development strategy as many people claim, wouldn't you expect them to immediately cease releasing open weight models and restrict access as soon as they start producing the frontier models? I am assuming this is what the USG thinks and is why they are trying to cut off the flow to foreign nationals ASAP.

LLMs are an undeniably valuable tool, and governments like to control those.

Your blog post doesn't get found by anyone in Google until you've built up your SEO mojo, your LinkedIn post isn't read without the followers you need to accumulate and your content has to get engagement for people to see it even then, you don't start off line with a million followers on X, etc.

I hate that this is true. It's the worst part about selling stuff online IMO and I found that you have to spend so much time doing it. In many cases, selling something online can be optimized to the extreme such that spend on marketing should be as high as possible and spend on the product R&D, manufacturing, support, etc should be minimized as much as possible. This equation gives you the most profit, but also gives the customer the absolute worst product that is possible to sell.

Capitalism doesn't really have a solution to this problem that I've seen yet.

That is, I would say that creativity requires that the new things generated be Evaluated. Without evaluation, and retention of the best, there is nothing created. The novelty flickers into existence but, if its value is unrecognized, it flickers away and is lost.

I really like the way he frames this here. I think a lot of people in the twitter comments (and maybe a few here) aren't reading past the introduction. He isn't saying AI systems are incapable of creativity and discovery. He is claiming generative AI without a harness is not capable of creativity and discovery. There needs to be some other system that "recognizes the value" of the novel idea and remembers it. He gives examples of where this value recognition step is automated and thus by his definition achieve creativity and discovery in a fully automated system.