I would assume so too, so the costs would not be so substantial to Anthropic.
HN user
cfcf14
I'm curious as to why 4.7 seems obsessed with avoiding any actions that could help the user create or enhance malware. The system prompts seem similar on the matter, so I wonder if this is an early attempt by Anthropic to use steering vector injection?
The malware paranoia is so strong that my company has had to temporarily block use of 4.7 on our IDE of choice, as the model was behaving in a concerningly unaligned way, as well as spending large amounts of token budget contemplating whether any particular code or task was related to malware development (we are a relatively boring financial services entity - the jokes write themselves).
In one case I actually encountered a situation where I felt that the model was deliberately failing execute a particular task, and when queried the tool output that it was trying to abide by directives about malware. I know that model introspection reporting is of poor quality and unreliable, but in this specific case I did not 'hint' it in any way. This feels qualitatively like Claude Golden Gate Bridge territory, hence my earlier contemplation on steering vectors. I've been many other people online complaining about the malware paranoia too, especially on reddit, so I don't think it's just me!
This makes me think it would be nice to see some kinda child of modern transformer architecture and neural ODEs. There was such interesting work a few years ago on how neural ode/pdes could be seen as a sort of continuous limit of layer depth. Maybe models could learn cool stuff if the embeddings were somehow dynamical model solutions or something.
The obvious next step here is to see how well this generalises to arbitrary inputs :)
Did your read the paper? Do you have specific criticisms of their problem statement, methodology, or results? There is a growing body of research indicating that in fact, there _is_ a taxonomy of 'hallucinations', that they might have different causes and representations, and that there are technical mitigations which have varying levels of effectiveness.
AI detectors do not work. I have spoken with many people who think that the particular writing style of commercial LLMs (ChatGPT, Gemini, Claude) is the result of some intrinsic characteristic of LLMs - either the data or the architecture. The belief is that this particular tone of 'voice' (chirpy sycophant), textual structure (bullet lists and verbosity), and vocab ('delve', et al) serves and and will continue to serve as an easy identifier of generated content.
Unfortunately, this is not the case. You can detect only the most obvious cases of the output from these tools. The distinctive presentation of these tools is a very intentional design choice - partly by the construction of the RLHF process, partly through the incentives given to and selection of human feedback agents, and in the case of Claude, partly through direct steering through SA (sparse autoencoder activation manipulation). This is done for mostly obvious reasons: it's inoffensive, 'seems' to be truth-y and informative (qualities selected for in the RLHF process), and doesn't ask much of the user. The models are also steered to avoid having a clear 'point of view', agenda, point-to-make, and on on, characteristics which tend to identify a human writer. They are steered away from highly persuasive behaviour, although there is evidence that they are extremely effective at writing this way (https://www.anthropic.com/news/measuring-model-persuasivenes...). The same arguments apply to spelling and grammar errors, and so on. These are design choices for public facing, commercial products with no particular audience.
An AI detector may be able to identify that a text has some of these properties in cases where they are exceptionally obvious, but fails in the general case. Worse still, students will begin to naturally write like these tools because they are continually exposed to text produced by them!
You can easily get an LLM to produce text in a variety of styles, some which are dissimilar to normal human writing entirely, such as unique ones which are the amalgamation of many different and discordant styles. You can get the models to produce highly coherent text which is indistinguishable from that of any individual person with any particular agenda and tone of voice that you want. You can get the models to produce text with varying cadence, with incredible cleverness of diction and structure, with intermittent errors and backtracking and _anything else you can imagine. It's not super easy to get the commercial products to do this, but trivial to get an open source model to behave this way. So you can guarantee that there are a million open source solutions for students and working professionals that will pop up to produce 'undetectable' AI output. This battle is lost, and there is no closing pandora's box. My earlier point about students slowly adopting the style of the commercial LLMs really frightens me in particular, because it is a shallow, pointless way of writing which demands little to no interaction with the text, tends to be devoid of questions or rhetorical devices, and in my opinion, makes us worse at thinking.
We need to search for new solutions and new approaches for education.
So uh, things are not looking so good for actual physics these days, I gather?
L-theanine (200mg) with around 100-150mg of caffeine has an extremely noticeable, positive effect on my ability to focus, feeling of "well-situatedness", and overall calmness. L-theanine by itself doesn't seem to do much. Caffeine on its own wakes me but makes me feel jittery and anxious, so it's definitely an interaction effect. Taurine has a much smaller effect on calmness, sans interactions - often indistinguishable from any other mild focus exercise like box breathing or stretching.
It's not - the FTC released a statement on this very topic a few months ago: https://www.ftc.gov/business-guidance/blog/2024/03/price-fix...
Police in large American cities are not likely to be of much assistance in this situation. Assuming they attend at all, I would expect them to not understand the nature of the issue and probably proceed to make it much worse.
After reading the paper, I'm really unsure what the novel contribution is. It feels like they're attempting to rebrand well-understood concepts within various fields (control systems theory, etc). The provided mathematical definition of antifragility is somewhat unconvincing too: it's not that it's wrong, per say, but in the effort to find something sufficiently broad to apply to many different fields of applied dynamical theory they've had to adopt a definition which is a bit unintuitive, and overly general.
This is really funny - it's bordering on truly absurd, almost incomprehensible madness to consider doing this seriously. I can't think of a single property you'd desire in a control system (state observability, auditability, guarantees on out-of-band input behaviour, stability under shocks, etc, etc) that would be present in an LLM control model.
I don't want to be disrespectful to the authors, and it's (vaguely) interesting to see how far they've been able to go with this, but this idea is still an abomination.
Was fun while it lasted! Will be interesting watching the internal story of the original lab unfold as it all becomes public eventually.
Yeah - more or less.
I still use it sometimes for trivial stuff like giving me recipe or travel inspiration (using the web search API), but I haven't been using it for any sort of algorithms/coding stuff.
It's useful for tasks which have a high tolerance for not being completely correct. Which is definitely a subset of all tasks that interest me, but it's a smaller subset that I originally thought it would be. It's just really hard to find out where the models are wrong when it's a complex question. If it wasn't, then I wouldn't need the tool, I guess.
Some strange claims in this post. The reddit post datasets are already 'out there' in the wild, and I'm fairly certain every other major LLM release has used their data. Also - did Midjourney "steal" DALLE-2's lunch? It's a restrictive service with essentially a discord-only based CLI.
Yeah, definitely. Combination of expert-system gating (some requests probably get routed to weaker models), distillation (for performance/cost), and RLHF lobotomization.
It's turtles all the way down, except for the final turtle, which is Fortran...
Lilian Weng's blog is my go-to example for an extremely high quality tech blog, it's truly remarkable how consistently excellent each post is. The only downside is the sadness I feel for being incapable of producing content even remotely near that level of quality myself.
This is 100% related to (suspected) fraud, anti money laundering, or other types of financial/political sanctions. You may be 100% innocent, but they will never disclose any information to you about their reasoning (and in fact it is illegal for them to do so). Sorry this has happened.
Absolutely not.
You have to prompt it correctly, non-instruction-aligned models don't behave like agent simulators by default.
Amazing post, agree with everything you've said. I've always felt that the problems with advanced MCMC methods (HMC, RM-MC, etc) are even more painful when one looks at approximate bayesian methods - ADVI (variational approximation), SGLD (langevin dynamics), and so on. My grad research was on ABC-SMC methods, probably the last resort of all last resorts.
There was a period a few years ago when it was all the rage to take arbitrary probabilistic programming models and just toss ADVI at it blindly with fancy tools like Pymc3 or stan. I feel like everybody eventually came to the conclusion that if the model was simple enough that you could guarantee that ADVI was actually correct, you didn't need it, and if it wasn't, you couldn't possibly verify you were approximating anything close to the true posterior.
At least with HMC it would generally explode when dealing with pathological geometry (multimodality, non-identifiability, whatever), whereas lots of the approximate methods will 'converge' to completely incorrect answers. I know there's been some work on determining if things have gone off the rails (https://arxiv.org/pdf/1802.02538.pdf), but I couldn't ever find a place where it felt safe to use this stuff blindly.
I wonder whether Bing has been tuned via RLHF to have this personality (over the boring one of ChatGPT); perhaps Microsoft felt it would drive engagement and hype.
Alternately - maybe this is the result of less RLHF. Maybe all large models will behave like this, and only by putting in extremely rigid guard rails and curtailing the output of the model can you prevent it from simulating/presenting as such deranged agents.
Another random thought: I suppose it's only a matter of time before somebody creates a GET endpoint that allows Bing to 'fetch' content and write data somewhere at the same time, allowing it to have a persistent memory, or something.
This was a really reasonable and interesting post by Stephen. I'm excited to see what the integration between an associative based model like GPT and a symbolic one like WA might bring.
It being from Schmidhuber's lab makes it dramatically more credible, in my views. They've been practically a decade ahead of everybody for ages now from a theoretical point of view.
He doesn't say this directly, but I suppose his comment here:
"The whole 'do an all-nighter to get the paper in the day before the dealine' is something that should have gone out the window after highschool. Not for kernel development."
is making an explicit comparison to the behaviour of young adults (namely: poor time management, crunching for deadlines by staying up late). I agree with you that it's a poor editorialisation and makes it seem like Linus is being harsher than he really is in his writing.
I discovered a while ago that you can ask GPT-3 for the output of even extremely obfuscated javascript code and it will produce the correct results most of the time.
Couple reasons come to mind: 1) The institutional apparatus required to maintain the monarch as head of state is reasonably complex and expensive, both from a legal point of view and in a financial sense. The governor general is an unelected official whose power technically exceeds any democratically selected person in our government, and they make almost $300k/yr to do essentially nothing but administrative ceremony. I know that in the big scheme of things, $300k (plus whatever other perks/bonuses they get, along with their staff) is just a drop in the bucket, but it feels a bit like theft from Canadians given this role doesn't need to exist at all. https://en.wikipedia.org/wiki/Governor_General_of_Canada
2) The royals do have a bit of a history of interfering in the politics of past colonies - the recent reveal a few years ago that Charles (I think...?) influenced the selection of high ranking Australian government officials in the 80's was a bit of a shock to many. Apologies, I can't find the right combination of phrases to yield anything in google except tabloid nonsense as a source. In any case, there is at least some evidence that the royals, either through wealth, old-school power politics, or direct action do play about with the politics of other countries, even when they explicitly should do not that. I detest this.
3) I really do not like what the royals stand for. Monarchism, colonialism, genocide, general crumminess. Spending £12 million to pay off one of the women Prince Andrew allegedly sexually assaulted as a minor, using public money, is an atrocious act and one I still cannot believe got so little media coverage here. Then there's all the general petty money issues: Prince Charles doesn't pay inheritance tax (why? Because, that's why). The royal family is except from all sorts of laws pertaining to wealth transfer actually, in some cases literally because it's coded into the laws themselves that Her Majesty is excluded. It's ridiculous. And here in the UK you still are not legally allowed to protest the monarchs. I mean - you can, but you can also be arrested for it. Just in the last 2 days there's been outcry as 2 people were arrested for anti-monarchist protest signs.
https://metro.co.uk/2022/09/11/woman-arrested-after-holding-... https://www.thenational.scot/news/21319718.protester-arreste... https://twitter.com/standardnews/status/1569270901685862403?...
It's time to let go of the past, try to build a future for ourselves that doesn't centre around the absolute worst parts of our collective histories.
As I Canadian I would strongly support removing the English monarchy as our head of state. I hope other countries do so as well.
- Unreliable and buggy core libraries (basic fp math, stats, autodiff) which seems to stem from an academic-style disinterest in focusing on the boring bits of language foundations
- 1 indexed arrays
- extremely slow start up times (1 min??)
- no way to create portable static binary apps (disadvantage vs rust/go)
- lack of tutorials, demos
- julia is not really faster for math when compared to jax/tf/pytorch/etc anyways