HN user

danbruc

9,533 karma
Posts7
Comments4,029
View on HN

[...] you can always reverse (a process) to determine [...]

There is a precondition - constant non-zero Jacobian - to the inverse existing and the inverse is claimed to be of a specific kind - polynomial. The counter example satisfies the precondition and by mapping two different inputs to the same output makes any inverse impossible, including polynomial ones. But maybe that is already ELI7.

The bet does not really matter, the central question is whether they have a fair coin or if they are trying to me in some way. Even without the magic store, I would be very suspicious of anyone approaching random people with an offer like that.

So I would certainly consider it likely, that they are trying to trick me. But the probability I would assign to this, would still be rooted in some frequency, somewhere under the hood I would try to estimate the possible situations leading to such an offer and in which fraction of them I will be tricked.

If I am doing a good job with that, then repeatedly being in this situation should result in me getting tricked with the probability I cooked up. If I am bad at figuring out the possible states and their probabilities, then I they will not match.

One could look at the likelihood of spontaneous coal fires and their effects, gather evidence about the activities surrounding the event, essentially trying to narrow down the set of possible states that would lead to the incident and then see what proportion of those states cause the sinking by fire, military action, or something else. But then attaching a number to that seems a dauting task.

Let us start with coin flips. You repeatedly flip a coin and the number of heads will come out to be about half the number of trials.

Where does that come from? It is not some intrinsic property of the coin, it comes from varying initial conditions. If you had enough precision when controlling your hand movements, you could in principle force an outcome with high probability.

But assuming you can not or at least do not do that, there is a certain set of initial states, some will lead to heads, some to tails, and each toss will start from a randomly selected initial state. So given my ignorance of the exact initial state, the coin will land heads with a probability equal to the number of initial states leading to heads divided by the number of initial states compatible with my observations of the initial state. [1]

Repeatedly tossing a coin will sample the set of initial states and the result will match the proportion of the number of states. At least as long as I am not wrong about the set of initial states.

The same applies to something like an election. I have imperfect knowledge about the state of the world but there is a set of states compatible with my knowledge about the world and certain subsets of them will lead to certain candidates to win.

[1] Maybe adjusted by some probability distribution over the initial states if they are not equally likely to be picked.

Something Bayesian. Despite my best effort I just do not get Bayesian probability, it more or less just does not make sense to me. Can you convince me otherwise? What is your best example of something with a probability that can not be analyzed in terms of frequencies or other proportions? And your Bayesian account of it must make sense, I am 90 % certain that P != NP and that is why I would take bets based on those odds does not cut it.

He will get the same as the others if he works the same number of hours to run the company. The dept would ideally be company debt and not personal dept, after all one reason to found a company is to separate personal and business risks. But if you invest personal money in the company, you get that money back plus some risk and opportunity premium just as you would pay back any other loan. What fraction that is will of course depend on the amount, premium, and payback period. Just investing some money will not provide you a perpetual income.

Envy is the completely wrong framing, it is about fairness. If you design widgets and 99 people build them and everyone works the same number of hours, then the fair share is 1 % for everyone. One percent of the work, one percent of the profit. If you take 2 % [1] while everyone else only gets 0.99 %, that is unfair. And look, the difference for each individual is tiny, 0.99 % or 1 %, that makes it easy to dismiss despite being the source of the unfair inequality.

[1] 1.99 % if we want to be precise.

Because the gap indicates unfairness and people care about that. If you are poor, raising the floor will be good enough, you primary worry will be sufficient income. Once you reach a comfortable life, your priority will - or at least may - shift, now you are in a position that allows you to worry if the situation is fair.

Also narrowing the gap is a way to raise the floor including in absence of growth while raising the floor without narrowing the gap depends on growth. But I think this is not too relevant in practice because taking a lot of money from a few wealthy people will often result in rather small changes as the money has to be spread among a lot of people.

I personally would not look for the way they reason in the weights, at least not directly. In principle I could replace a large language model with a map from all possible input strings to output token or output token distribution without any weights. I have a hard time imagining how you would even tell, at the level of weights and activations, if the next token being the is the result of some proper reasoning or a hallucination. But those weights do not exist for the sake of it, they encode a lot of text the model has seen during training, and I would imagine this is what drives the reasoning. Can you evaluate the following polynomial ... will be related to To evaluate a polynomial ... seen in the training data. This is the level at which I would look for the reasoning, memorized patterns how to do specific things, maybe with some kind of placeholder variables for generalization. Ultimately such a structure would of course also be represented in the weights but I could imagine that this makes it unnecessary hard to understand. Or maybe not, maybe the learned patterns are so complex that they do not have a simple representation.

AI 2040: Plan A 11 days ago

I remember thinking exactly the same thing around 5-7 years ago, in the GPT-2/GPT-3 era.

ChatGPT was announced three and a half years ago, 30-11-2022.

What do you prefer, a plane crash or a nuclear war? But that does not matter, I did not say a airplane crashing is preferable over something, I said it is not a bad event for everyone. It is bad for the people on board and their friends and families but there are also people that benefit from a crash. That still does not mean that people benefiting from the event are in favor of more airplanes crashing. And for the majority of people it will essentially be neutral, half the planet might never even hear about this particular crash and for many more it will be a thirty second news segment they once saw.

The point of the example is that something that would generally be judged as obviously bad often turns out to not be bad from all perspectives. The same obviously also holds for good things. But I did not intend to imply any economics. The crash investigation might uncover an issue that gets fixed and prevents future crashes, also in this sense the crash had a positive effect. The better outcome in some overall sense would probably still have been for the airplane not to crash and the families spending the money on something better than funerals.

[...] some detractions are more worthy of consideration than others.

This is probably quite hard to justify in general.

I want zero cents of my money to be wasted on climate change, I live now and I want to enjoy my life as much as possible. I do not care if the planet gets burned to a crisp in a hundred years when I am dead.

One can certainly have a different perspective and vehemently disagree, especially if you have children that will have to live through that future, but otherwise? How would you argue that this position is somehow less valid than any other?

1. An airplane crashes, everyone dies. Clearly a bad thing. Not so quick. For the funeral industry this means additional business, a good thing. And this is true for a LOT [1] of things, they are not good or bad, right or wrong, their judgment depends on perspective and personal preferences.

2. Which means that there is generally no policy that makes everyone happy. So you need a party with a program that aims at finding compromises that are acceptable for everyone.

3. But nobody will vote for such a party. Why would you vote for a party that gives you 50 % of what you want if there is a different party that is more aligned with your views and preferences and promises to give you 90 % of what you want?

4. In consequence the political direction tends to hop between extremes instead of settling on compromises. One group gets really unhappy with the current situation, shows up for elections, votes their party into power, moves the situation into the direction of a different extreme, until others get unhappy enough to start the process all over again.

5. Even in political systems where [sometimes] a coalition of parties exercises the power and they are forced to compromise, the outcome is all but ideal. Things move slowly because finding compromises is hard if you do not really want to compromise. Voters look down on the party they voted for because they are not delivering what they promised but only compromises.

I guess the moral of the story is that the voters have to realize that their view is not the only valid one and that voting for compromises would probably yield better outcomes than voting for extremes and either going in that direction for some time until turning around or maybe arriving at a forced compromise that no one voted for.

[1] Exercise for the reader, find something politically relevant that does not depend on perspective and personal preferences.

This already existed ten years ago in the desktop version, not sure if it also was in the web version all the time.

CrankGPT 1 month ago

Not only nuclear power plants have cooling towers. Here [1] is, for example, a coal-fired one in Poland.

[1] https://en.wikipedia.org/wiki/Coal-fired_power_station#/medi...

EDIT: To elaborate a bit, if you are burning oil or gas in a turbine, you do not need a cooling tower, the waste heat goes into the atmosphere with the exhaust. If you use fossil or nuclear fuels to produce steam for a steam turbine, you either need a river with enough flow to not boil all the fish if you reject the waste heat into it or you need a cooling tower to reject the heat into the atmosphere.

I am not a mathematician and did not read the unit distance solution too carefully, but my impression was that it used a variation of a known technique to solve the problem. And that makes perfect sense to me, there are a lot of techniques and lot of less relevant problems, I am not surprised that one can solve some of them with known techniques that just nobody has tried [hard enough] before. I am much more sceptical when it come to the important unsolved problems where every known technique has probably been tried several times over. In those instances it will probably take a true leap in understanding to solve them and I am sceptical that large language models are well suited for that because of the way they work.

This is just looking at it from a different perspective. Both, one 64 bit integer and the product of two 32 bit integers represent a number up to 2^64 with 64 bits. But while all 64 bit integers are unique, there are, as you say, several representation for some numbers as the product of two 32 bit integers and therefore it is impossible to represent all 64 bit integers. Commutativity alone costs you about 50 percent of all numbers in the range as x * y and y * x represent the same 64 bit number but with two different representation as a product of two 32 bit numbers, at least if x and y are different. But this tells you nothing about the numbers that you can not represent, only that they must exist. I was looking at it from this other perspective, which numbers are not representable as a product and why.

Each x is prime with probability 1/ln(x), each x has M/x multiples less than M, as a fraction of M that is just 1/x. Together that makes 1/(x ln(x)) with the indefinite integral ln(ln(x)). If we plug in 2^32 and 2^64 [1], we get ln(2). So about 69.3 % of all 64 bit integers should have a prime factor larger than 2^32 and therefore not be the product of two 32 bit integers. That leaves about 13 % unaccounted. Three prime factors all larger than 2^32/2, five prime factors all larger than 2^32/3, and so on cannot be packed into two 32 bit integers. Not sure to how much this will add up.

[1] The bounds are important because they guarantee that there is at most one prime factor from that range and this ensures that we are not double counting anything. If the upper bound was larger than the square of the lower bound, then we would have to worry about double counting numbers with more than one large prime factor.

No RAM. Instead of having a general purpose multiplier that multiplies an input with a weight stored in RAM, just have a multiplier that hardcodes the weight. In some sense replace each weight with a specialized multiplier and wire them together with accumulators and activation functions in between. And some registers for pipelining. If one goes for four bit quantization, one could have sixteen optimized multipliers, one for each possible weight, and the one just selects and connects them according to the model weights and structure.

Example. If you have a neuron with 16 inputs each 8 bit wide and with a 4 bit weight per input, you will have 16 specialized multipliers each scaling its input by the corresponding weight and then the 16 scaled inputs feed into an adder tree and finally an activation function.

Did some try to estimates what it would take to bake interference for a capable large language model into silicon so that one can pipeline inputs through it and produce outputs at one token per clock cycle?