Can you elaborate on running them "in conjunction"... are you running the same query on multiple models and then using a third model to judge or make consensus? or am I misunderstanding completely. I'd like to understand how these small models "run together"
HN user
ethanpil
I opened the tab in the background and when i got around to it later somehow i lost the game.... i think you need an official "start" button to prevent this.
Another recent discovery of mine in this genre: https://news.ycombinator.com/item?id=48881881
very cool thanks for sharing.
Wow. As a comparison, I just opened a new Google Maps tab in Chrome. According to the Chrome Task Manager, the tab alone uses 433mb RAM and 34mb GPU memory footprint after first load.
I'd like to study your setup. Would you be willing to share? Perhaps a github repo of your 5 extensions or even a pastebin if you would be so inclined. I would be grateful to learn more about this by studying from your success...
Another hot take from him in 2018 is "Many people I respect here, lately identified Facebook as the root of all the evil. I want to start this thread about why I disagree..."
Per the "Availability" section of the page, seems like should come back to all plans eventually...
* From today through June 22, Fable 5 is included on Pro, Max, Team, and seat-based Enterprise plans at no extra cost.
* On June 23, we’ll remove Fable 5 from those plans. Using it after that will require usage credits. If capacity allows, we’ll extend the included window.
* After this point—when sufficient capacity allows us to do so—we aim to restore Fable 5 as a standard part of subscription plans. We intend to do this as quickly as we can.
What's Google's business case for releasing open models? Don't get me wrong, I am grateful and appreciative of these releases. I'm trying to understand how it fits into their bigger picture as a for profit company? Are they not helping competitors build on the novel technology they have developed?
Is it simply goodwill and/or marketing? Or am I missing something strategic?
Anyone here remember the early days of WhatsApp, pre-Facebook, when it required an annual subscription fee of $1?
Why not state that?
The table comparing eval scores shows the following:
Agentic Terminal Coding (Terminal-Bench 2.1) Opus 4.8 74.6% GPT 5.5 78.2%
Then, when you scroll all the way down to the bottom Footnotes section it says
"Terminal-Bench 2.1: We reported scores for all models using the Terminus-2 public harness. GPT-5.5’s reported score with the Codex CLI harness is 83.4%."
Can you share the GGUF for this specific success story? I'd like to try it for myself.
> It took two initial prompts and a few tiny follow-ups. GPT-5.2 running in Codex CLI ran uninterrupted for several hours, burned through 1,464,295 input tokens, 97,122,176 cached input tokens and 625,563 output tokens and ended up producing 9,000 lines of fully tested JavaScript across 43 commits.
Using a random LLM cost calculator, this amounts to $28.31... pretty reasonable for functional output.I am now confident that within 5-10 years (most/all?) junior & mid and many senior dev positions are going to drop out enormously.
Source: https://www.llm-prices.com/#it=1464295&cit=97123000&ot=62556...
What's the use case for this?
Reading this document I can now confirm 100% that at least 1 AI has Em Dashes embedded within its soul.
Another interesting approach IMHO is https://github.com/gnat/surreal
How do you know this happened? I thought it was an abandoned project until I saw this post. I've been diligently checking weekly for new releases but nothing for almost a year...
Had a similar issue - wanted to get all the files from the response without too much work, so I opened a new tab and vibe coded this in about 4 minutes. Tested it on exactly 1 case: a previous Sonnet 4.5 response, and worked well.
This is wonderful i'll be using this a lot. Would be great to see a filter for Tiny vs Mini as well as CPU. AMD/Intel and i3 i7 i9, etc, maybe even generation, etc
Looks like a clone or fork of the UnRaid docker interface
Why are they getting worse?
You have those AHKs somewhere online? Would love to peruse them...
With all of these great improvements in recent months i'd love to see an updated roadmap... https://github.com/containers/podman-desktop/wiki/Roadmap
What would it take to get this on modern (or even last generation) phone hardware?
I bet we could do everything we want with tremendous simplicity and out of this world battery life... Probably would make a PinePhone feel like a Rolls Royce.
Nice. My next step: Figure out how to make a web extension 1 click button. Tab to Monolith to Joplin with a tag.
Good luck catching a layover if you are Jewish or have an Israeli stamp on your passport. Last time I checked the staff of Saudi destined airlines wouldn't even let "undesirables" on an incoming flight...
I have been curious about using computer vision to bulk scan and identify valuable coins. Have systems like this been put into use already? If not do you think they might dilute the value of "rare" coins by creating many more "finds" then were typically possible without the technology?
Remember when Google was cool and not evil and released their book scanner project for free? https://code.google.com/archive/p/linear-book-scanner/
"Google hereby grants to you a perpetual, worldwide, non-exclusive, no-charge, royalty-free, irrevocable (except as stated in this section) patent license to make, have made, use, offer to sell, sell, import, transfer, and otherwise run, modify and propagate this design..."Also Rudy Kurniawan who sold counterfeit wine and was convicted about 10 years ago. A quick search turned up this article... didnt read it. https://thehustle.co/the-man-who-sold-millions-in-counterfei...