HN user

joshred

101 karma
Posts0
Comments48
View on HN
No posts found.

I don't think these are the same. Outlawing CSAM gives law enforcement the ability to shutdown markets and prevent commercial distribution of CSAM. Sexually abusing children is heinous, but sexually abusing children for financial gain is even worse.

From what I've read, that's already part of their training. They are scored based on each step of their reasoning and not just their solution. I don't know if it's still the case, but for the early reasoning models, the "reasoning" output was more of a GUI feature to entertain the user than an actual explanation of the steps being followed.

Gerrymandering already exists. Voter suppression was huge in the past, and may become huge again. The supreme court made sure of that.

And also... the supreme court keeps issuing partisan decisions.

So... what is left? Number 3?

I guess you're arguing that federalism protects people, but how does it do that in a way that isn't already being eroded?

I work for state government. We've used the ACS survey to try and determine whether we were unfairly targeting non-native English speakers with some of our decisions. It's also used a lot in academia.

If I had to guess, commercial organizations have access to more invasive and higher quality data that they obtain through credit card companies, lexus-nexus or other data brokers. This attitude mostly harms organizations involved in the social sciences.

It sounds like they are describing a regex filter being applied to the model's beam search. LLMs generate the most probable words, but they are frequently tracking several candidate phrases at a time and revising their combined probability. It lets them self correct if a high probability word leads to a low probability phrase.

I think they are saying that if highest probability phrase fails the regex, the LLM is able to substitute the next most likely candidate.

They might not now how whisper works. I suspect that the answer to their question is 'yes' and the reason they can't find a straightforward answer through your project is that the answer is so obvious to you that it's hardly worth documenting.

Whisper for transcription tries to transform audio data into LLM output. The transcripts generally have proper casing, punctuation and can usually stick to a specific domain based on the surrounding context.

This is the high-level explanation of the simplest diffusion architecture. The model trains by taking an image and iteratively adding noise to the image until there is only noise. Then they take that sequence of noisier and noisier images and they reverse it. The result is that they start with only noise, and they predict the removal of noise at step until they get to the final step (which should be the original image (or training input)).

That process means they may require a hundred or more training iterations on a single image. I haven't digested the paper, but it sounds like they are proposing something conceptually similar to skip layers (but significantly more involved).

I think they're fantastic at generating the sort of thing I don't like writing out. For example, a dictionary mapping state names to their abbreviations, or extracting a data dictionary from a pdf so that I can include it with my documentation.

Welcome, ACLU 9 years ago

Cultural conservatives are authoritarian. They argue for restricted social liberty.