Same here, as we're sitting in the middle between requests and what budget constraints are allowed given a particular token allowance there can be 10 ~ 100 milliseconds improvement in the UX (TTFT) given such massive tokenization speed up.
HN user
flockonus
Building tokenbank.ing
Ah yes, i missed it because i cmd+F for "pric" & "cost" - but now i read more carefully, it makes sense. It also speaks about the biz model.
Personally i wouldn't change a thing, other than perhaps add one of those words (if at all).
At first i found odd that pricing is nowhere in the website, but then i found it burried under FAQ (free!) - is it to weed out people who are focusing on free rather than going deeper into the experience?
How much does attending the Recurse Center cost? RC is free for everyone. You will never receive a bill from RC.
How can you afford to make RC free? RC has a built in recruiting agency. Companies pay to hire RC alumni. This payment never comes out of your salary.
There is also the 1bit version @ 3.9 GB that retains 90% of the intelligence - quite a feat!
Excited to read if you ever get to publish! In particular 1. made me even more curious.
I do get most of my AI value on prototyping [small codebase] & summarizing.
Cool! I also remember from decades ago a friend in high school taught me "SOS" and still stuck with me: 3 short, 3 long 3 short, i suppose ... - - - ...
That was a departure from how I normally work since I did not write a single line of code. I will likely write something about my A.I framework (and opensource it) next month.
If the author is around, super curious how they got to enjoy their workflow here in a side project, working with AI. This kind of situation, to me, is often where the joy is gone.
DOGE found an actually highly efficient Federal government
Wish we could see the evidence for that.
It would be true if there was a unified "Anthropic" entity making every decision from pure rationality. Instead, more tokens increase Claude Code team's metrics of token usage, which most likely has a KPI around token usage and adoption.
To remind Goodhart's law: "When a measure becomes a target, it ceases to be a good measure".
..also to parent's point, yes the upsell is only appealing once user run's out of tokens.
I suppose it's like gradient descent, often times you get stuck in local minima and there is no inertial force to overcome the next peak.
Curious for what an MTP only result would look like, both in terms of output quality & tk/s ?!
Looking at this gave me a renewed (even if brief) sense of appreciation for our society.
Even with all of its problems, for this moment in time our society is operating effectively enough that humans can engineer interesting, beautiful, culturally rich landscapes (cities) in quite a significant size scale.
Cheers to us, in 2026.
Some irony in so many posts about AI becoming more capable at programming, at the same time, top post on hackernews is a game about where you code by reading a magazine like it's 1997.
not so interesting to compare to
Absolutely disagree here, something that is considered good practice is very interesting to compare to!
Oh while coding i sure think far more "tokens" than i output to the editor, at least 4->1.
Now, if those should be counted in the process or only the output is a harder question.
Interesting point, i didn't consider the thought process as tokens.
For coding tasks 27B is reported to be much more effective, altho you can probably only run 4b or 5b quants @ this memory.
Recommend https://www.reddit.com/r/LocalLLaMA/ as a great source for this type of discussion.
Curious about the other way around, how many tokens per second a productive developer codes in a day?
Curious to see if to what intensity the Moon will "ring like a bell" at this one.
ref: https://books.google.ie/books?id=6QAAAAAAMBAJ&pg=PA56&lpg=PA...
How this will age:
I do not and will not use the internet, in any form, for any purpose.
How to check if your voice is being misused
I love that the answer here is basically.. - you don't -
But maybe mitigate at unreasonable personal costs.
How about services simply stop taking public information as proof of identity?
That may be true for OpenAI, less so for Antropic - which has much better margins. Both of these companies CEOs have come in public saying the same.
No doubt as of currently Google has a better business. But the same argument could have been said about Instagram or Whatsapp before Facebook (now Meta) acquired them.
The bigger the [dense] models the more inference tends to take, it seems pretty linear.
In that sense, how long you'd need to wait to get say ~20tk/s .. maybe never.
(save a significant firmware update / translation layer)
I know where my answer lies in that; but i don't claim to be an objective truth.
For example OpenAI has been caught sharing data with the gov. agencies.
While they do make this argument, realistically anyone sending their prompt/data to an external server should assume there will be some level of retention.
And more so in particular, anyone using Darkbloom with commercial intents should only really send non-sensitive data (no tokens, customer data, ...) I'd say only classification tasks, imagine generation, etc.
My motivation was quite different, and i'd like to encourage more people to consider the same.
Often times narcissistic power grabbing (often technically incompetent) engineers become managers, like it was the case a previous team I've worked at and it was quite penalizing to the whole team.
I've realized that either i can be the one managing and try to do good, or be at the mercy of another manager; chose the first.
The readme seems very unclear about what it does. Anyone has a practical example of it?
nit: the menu on your portifolio feels drunk.. the last thing i want as user is links skidding away as i try to click them
This post says remote, but then Greenhouse for the manager says it requires to be in NYC.
Would a Canadian candidate be considered?
I was previously manager at several monetization teams at Asana.
Anyone who has a mobile phone has been tracked by their phone provider forever, with the accuracy of a couple blocks. Smartphones only bring more trackers to the equation in the form of apps.
What's the material concern to tracking that glasses add?