You don't know what you are talking about. Obviously refusal circuitry does not live in one layer, but the repo is built on a paper with sound foundations from an Anthropic scholar working with a DeepMind interpretability mentor: https://scholar.google.com/citations?view_op=view_citation&h...
HN user
robertk
https://robertk.at.hn
[ my public key: https://keybase.io/robertzk; my proof: https://keybase.io/robertzk/sigs/MMuXENMT-5k_cTFWUS71BLbqSaYKTQhCxNURSaOMgkQ ]
Why not just open it inside of and print to a static image output within a fully sandboxed Docker container?
Why not leak a dataset of N full text paraphrasings of the material, together with a zero-knowledge proof of how to take one of the paraphrasings and specifically "adjust" it to the real document (revealed in private to trusted asking parties)? Then the leaker can prove they released "at least the one true leak" without incriminating themselves. There is a cryptographic solution to this issue.
It’s slightly biased. ( P(even) = 0.5702; Bias = +0.0702 (about 7 percentage points toward heads) ). You can use this Claude Code prompt to determine how much:
Use your web search tool call. Fetch a list of English words and find their incident frequency in common text (as a proxy for likelihood of someone knowing or thinking of the word on the fly). Take all words 10 characters or longer. Consider their parity (even number of letters or odd). What is the likelihood a coin comes up heads if and only if a word is even when sampled by incidence rate? You can compute this by grouping even and odd words, and summing up their respective incident rates in numerator and denominator. Report back how biased away this is from 0.5. Then do the same for words at least 9 characters to avoid “even start bias” given slight Zipf distribution statistics by word length. Average the two for a “fair sample” of the bias. Then run a bootstrap estimator with random choice of “at least N chars” (8 <= N <= 15) and random subsets of the dictionary (say 50% of words or whatever makes statistical sense). Report back the estimate of the bias with confidence interval (multiple bootstrap methods). How biased is this method from exactly random bits (0.5 prob heads/tails) at various confidence intervals?
“Slightly fringe”
I am sorry for your loss, Aella. I sobbed with you.
“Each passing minute is a greater percentage of the final minutes we have,” and yet “these [final] seconds are so soft”.
Death needs to die, some future dying day, not yet.
from everyone who’s had a mom, we join you: “Momma, I love you”.
You may be interested in: https://www.anthropic.com/research/sleeper-agents-training-d... https://arxiv.org/abs/2404.13660
The Apple paper does not look at its own data — the model outputs become short past some thresholds because the models reflectively realize they do not have the context to respond in the steps as requested, and suggest a Python program instead, just as a human would. One of the penalized environments is proven impossible to solve in the literature for n>6, seemingly unaware to the authors. I consider this and more the definitive rebuttal of the sloppiness of the paper: https://www.alignmentforum.org/posts/5uw26uDdFbFQgKzih/bewar...
If I read a comment that has any probability of changing my mind about a fact or opinion, I always go to the user page to check their registration date. No hard cut-off date but I usually discount or ignore any account >= 2020.
Shawn, there is a mildly redacted version available at https://huggingface.co/datasets/monology/pile-uncopyrighted
No, it doesn’t. This concerns a corporation subject to legitimate national security concerns, not “a person, or a group of people.”
Very cool result but the title is overselling the "AI" contribution. It seems like they trained a few standard binary classifiers (Naive Bayes, decision trees, kNN). The novelty is the independent variable coming from an attribute precomputed for many known elliptic curves in the LMFDB database, namely the Dirichlet coefficients of the associated L-function; and the dependent variable being whether or not the elliptic curve has complex multiplication (CM), an important theoretical property for which lots of flashy theorems begin with assuming whether or not the curve has CM. They go on to train another binary classifier (and a separate size k classifier) to determine a curve's Sato-Tate identity component using the Euler coefficients and group-theoretic information about the Sato-Tate group (constructed by randomly sampling elements and representing the two non-trivial coefficients of their characteristic polynomials as independent variables in the classifier). They also run a PCA: https://arxiv.org/pdf/2010.01213.pdf
The cool part is that they then stepped back and scratched their heads wondering why the classifier was so good at achieving separation for these dependent variables in the first place, and plotting the points showed them to be (non-linearly) separable due to a visually clear pattern! The punchline and the reason it's so important to understand these data points, the Euler coefficients for elliptic curves, is because they contain all the relevant number-theoretic information about the curve. With some major handwaving, understanding them perfectly would lead to things like the Langlands program (and some analogues of the Riemann hypothesis) getting resolved. These wide reaching conjectures are ultimately structural assertions about L-functions, and L-functions are uniquely specified by their Euler coefficients (the a_p term in their Euler factors). Will murmurations help with that? Who knows, but the more patterns the better for forming precise conjectures.
Relevant intersectional credentials: I have lead ML engineering teams in industry and also did my doctorate work in this area of math, including using the LMFDB database referenced in the article for my research (which was much smaller back then and has grown a lot, so very neat to see it's still a force for empirical findings!).
Costco only takes cash and debit, not credit.
Not really. Only a tiny slice of the historical person’s memories and persona is recorded. There is a lot more entropy to their representation that died when their brain did. Ergo, whatever “perfect” simulacrum is presented will need to infer the gaps and ultimately be fictional.
By the pigeonhole principle, there is a sentence that writes out its entire SHA256 representation this way. Alternatively, the map from these kinds of sentences with 256 terms to 2^256 given by SHA256 admits a fixed point.
Just a note that this is by Geoff Anders from Leverage Research, an organization historically plagued with some controversy in the level of psychological experimentation it is willing to perform on its members:
https://medium.com/@zoecurzi/my-experience-with-leverage-res...
My best friend’s wife is in the process of dying right now after qualifying for and self-choosing hospice following persistent and progressing medical issues. He was looking for graveyards yesterday while she continues to pass… I am one of his primary support structures and this is hard for me, too. I just want to be as normal as possible for/around him, to be a rock. But I have never been in this position for someone before and I don’t know what would be most helpful. If anyone has, or possessed the empathy and EQ to be truly attuned to an impossible situation like this, can you please reach me through my profile or respond to this comment? With gratitude in advance.
I am someone for whom the two universes are the same. I freely choose to act exactly in accordance with how I would act if I were steered by neurons and atoms, and neurons and atoms alone. And that is also the truth, so I am perfectly happy. The universe with and without free will, to me, is metaphysically identical - because of how I have chosen to live and be happy.
The statement “free will does not exist” does not imply the world isn’t Minecraft. It is. Arbitrary actions allowed by physics are indeed allowed. The problem is unpacking that you can do what you want. It is the act and nature of wanting that is also fully constrained by physics. There is no “you” and no “want,” no notion of choice that isn’t simply isomorphic to doing the basic math of quantum mechanics and molecular biology… you are “merely” math executing in time and space, and you are and shall never be nothing more. https://secularsolstice.github.io/speeches/gen/Nothing_Is_Me...
What you do and think is exclusively dictated by the input of information going into your brain coming from perception encoded as electrons and ions flowing through the atoms consisting of your neurons. If someone had a printout of the small sub-manifold of spacetime that is you they could read out everything you ever thought and did and it would all make sense (meaning, follow from Schrodinger’s equation). You have no choice and you cannot defy physics. Even if you run into the street, it was so decreed, because you were deterministically ordained to be contemplating free will and attempt to violate it with a “random” action, which in conjunction with the random seed you have chosen (the current state of your neural ion flows) led you to deterministically run into the street. With, as always, a perfect causal pathway and explanation — no matter how free and random it seems. What you do next, always, is decided exclusively by what in the universe happened just before, and no other auxiliary or theological inputs.
Regret, like everything else, is an evolutionary corollary built out of destiny optimization that allows for counterfactual consideration of painful paths to avoid similar outcomes in the future. In the end, what physics decrees shall occur shall occur and the only thing we can do in those times of war is burn down the ships behind us and continue charging forward in this strange yet charming universe we find ourselves shackled within.
In the end, we must imagine Sisyphus—the final mirror—happy.
Dude, spaceman_2020, you are young a.f. still. There are predicaments and Chinese finger traps much much worse still. Here is your quote of the day:
https://www.goodreads.com/quotes/6495956-for-what-it-s-worth...
Nothing. Free will doesn’t exist so I just accept what comes my way and naively trust the universe not to be too much of a d$&#. Hopefully that gets me to the finish line.
https://www.financialsamurai.com/best-time-to-start-a-busine...
“ The Best Time To Start A Business Is During A Downturn”
Edge tails of high variance inputs lead to very different values of the computational outputs of thought derived from those inputs, news at 11.
In theory I agree with you. Unfortunately this is a very, um, privileged way of viewing this aspect of the world. A lifetime of selection bias clouds, wherefrom this strategy works for you, but would not if tried by an other.
This seems like a great use case for those long zoom meeting codes.
I am not the OP and I’m not sure how to directly message you - there is no contact info on your profile and HN does not have DM support. Nevertheless, I am very interested in your last sentence - may I take that as well? Happy to provide more context over the email listed in my profile. Thank you.
Hey Syzygies. Can you please reach out to me at my HN bio? I am conducting an analysis of mathematicians and Fields medalists’ proclivity to remain within pure mathematics terrain vs context switch to some worldly contribution, particularly in light of some of the pressing problems of modernity. You seem to have insight in this subject. If you have the few minutes of time, I would appreciate you reaching out. I can compensate for survey time as well. Thank you in advance!