No, balanced ternary, for example, uses {-1, 0, 1}. The system you're discussing is balanced quinary (base 5).
HN user
numeri
You could add a toggle, so that if someone's happy to wait for the key setup, they can try the full end-to-end process
I read the comment you're replying to as saying, "in the US, but other countries may have different policies that result in lower recidivism, and that might change the conclusion; maybe people aren't inherently criminally insane, but can become useful members of society, if given a chance"
Seems to echo (but in a watered down form) many of the ideas in https://gwern.net/guardian-angel, which gave me a lot to think about last week
I've not written up anything, no. I think I'd have a hard time doing so without just feeling like I'm bragging about myself, which I don't like.
There's still a definite gap between me and native speakers, that shows itself primarily in the effort required, but I'm definitely near native (pass as German in all social settings, although an hour long conversation will usually tease it out due to my unfamiliar first name or small-talk topics, rarely but occasionally due to mistakes).
I prepared by doing two practice exams and about 5 filmed and timed practice presentations, and that was over preparing for me. Experiences vary, and I do think I'm a bit towards the outlier side, but it's left me convinced that the whole "native speakers might not pass C2" thing is overblown.
Most Germans won't be able to pass a C2 test
That's not true, but it is a commonly shared myth. I've taken and passed C2 with the highest mark in every category (I moved here when I was a young teen, wanted to know if I would pass it after hearing years of people saying things like you're saying).
Most Germans would easily pass C2, although I think they'd have to be well-read/possibly university educated to get high scores (mostly need to be able to read quickly, give a semi-structured presentation and write a persuasive essay).
For what it's worth, I could run linguistic laps around all the other test takers there that day, and I assume at least some of them passed.
No, quantization is applied to model weights or the KV cache (the model activations of all past tokens), and is just storing everything with lower precision (carefully, so that it doesn't hurt performance much).
Sending an image of text instead of text reduces the number of input tokens, but they're still being processed by the model at the same precision. This probably also hurts performance in some way – the question is by how much.
Deep seek OCR is an LLM, just one trained/post-trained specifically for OCR.
Exact details of text to image compression ratios are of course extremely dependent on the model architecture, training data, training objectives, etc., so there's probably not too much justification for generalizing to all models
Are you writing general use programs in it, then? Have any good examples?
A future system that works like you described would be awesome. It'd be like community-sourced peer review (although by community I mean a community of experts in different fields, not arbitrary individuals).
I'd love a statistician's review on a ton of the papers I read.
The problem is that what people care about are the "black swan" causes of death, i.e., the cases the actuarial table is wrong.
Prices for training have dropped immensely in terms of research required, code efficiency, algorithmic/sample efficiency, and possibly also hardware (I'm not qualified to say without looking it FLOPS/dollar, or even to be certain that's the right metric here).
I mean, it might listen to him. We have no clue, which is the problem.
There's a large gap between making up words and an actually native text distribution. LLMs have a clear pattern, clear tells, a "feel" in English, and it's normally even more pronounced in non-English languages.
Lots of bias towards English sentence structure, idioms, etiquette, etc.
One context I could imagine is a young person with shaky grasp of English trying to come up with an interesting school/university project via conversations with an LLM set up as an OpenClaw agent.
It's got the right combinations of inexperience, cluelessness, panic, expectations that Westerners are rich, and hopes of others being willing to fix their mistake.
especially because this is the most painfully glaring flaw in their plan. Their solution is for an inference provider to... store the KV cache (which they can compute!) on-premise, on their own disks, but pay some third party for it?
I've had it happen. I ran an experiment, taking a couple hours and producing ~2 GiB of files. One of the results looked good, so I told Claude Opus 4.5 (at the time) to commit the code changes, upload the important file to cloud storage, then clean up the rest.
I then saw it run `rm -r results/`, before messaging me: "Now all that's left is for you to upload the successful results, then I'll delete the rest!"
Why did it not upload the files itself, when it had been using the cloud storage CLI during that session? No clue. I do accept that I could have and should have just uploaded the file myself. It would have taken 3 seconds to type.
To be fair, it is good to know that it disobeys simple instructions like "don't examine my git history" far more than other models. (It should of course be a different benchmark, so as not to conflate things.)
It's not a great sign for alignment.
I would just warn that you may not be able to recognize what is worth learning at your stage.
Intuition for library design and the architecture of software packages/external APIs is something you can only learn by doing.
I have DSPD as well, and was pleasantly surprised to see how much of the article discussed DSPD.
That being said, I do think a lot of what the author is saying flies right in the face of traditional advice, esp. the suggestion that we should all just free-sleep and rotate around the clock. I personally find myself happiest when I'm entrained to the 24-hour cycle, but at my own natural offset. Whenever I've been cycling the day it's felt miserable, uncontrollable and exhausting.
To be fair, the author did claim that you can fully solve this by completely cutting out after-dark electronics, but I've tried pretty intensely to do exactly that for extended periods in the past, and didn't see any progress. I do sleep amazingly when camping, though, and the delay is lesser than normal (still definitely there).
11/20 for qwen/qwen3.5-flash-02-23 in Claude Code, with effort set to low.
No, that's what the headline implies, and the body of the article doesn't support at all. It's (currently, and with no indication of intent to change this) two separate branches of their business.
but Taalas had to quantize Llama 3.1 8B to death to get it to fit. It can't produce coherent non-English text at all.
and if I was to guess, the latest generation of models (Claude Opus 4.6, GPT-5.3-codex, etc.) differ from Opus 4.5, GPT 5.2 primarily in the addition of deeper, more difficult (most likely agentic and coding-based, like Terminal Bench) tasks to their RLVR training.
I could be completely off, as my intuition here is fully based on public research papers, but it seems to explain the current state of things fairly well.
No, Python or units[1] is always a better choice if I'm near a computer (and I nearly always am these days, unfortunately, I suppose). I do have three wonderful slide rules, though.
Introducing a solid zero-knowledge age verification option is the opposite direction of ending anonymity in the Internet, which other parts of the same governments are also working on.
So yeah, I'll gladly trust and cheer on the part working in the right direction.
I'll just throw in support for gaming on Linux – it's pretty nice feeling these days! I still have the occasional (once every 5–8 months?) update cause a short-lived bug, but it's a very justifiable trade-off to avoid Windows these days.
This is written by someone who's not an AI researcher, working with tiny models on toy datasets. It's at the level of a motivated undergraduate student in their first NLP course, but not much more.
One sign would be occasionally changing course in response to overwhelming employee feedback. If that never or almost never happens, the feedback is being ignored, not taken constructively and not followed.
This isn't right – calibration (informally, the degree to which certainty in the model's logits correlates with its chance of getting an answer correct) is well studied in LLMs of all sizes. LLMs are not (generally) well calibrated.