HN user

DoctorOetker

1,797 karma
Posts8
Comments2,268
View on HN
Qwen 3.8 3 days ago

there is a very good reason to believe the parameters are still highly redundant: just as one example, recently there was research into repeating carefully selected blocks of middle layers, and seeing improvement, it turns out there are 3 types of layers: the initial layers that translate from natural language tokens to some kind of LLM-specific universal "thought space", reasoning blocks of layers that can be repeated operating in "thought space", and then a final stack of layers for translating from "thought space" back into natural language space.

lets ignore any compressibility in these initial and final layers which recognize lanuague, jargon, parsing natural language to "thought space" or back, instead let us look at the repeatable blocks, if inserting extra copys of stacks of layers only improves the result, its as if such correctly scoped middle layers look at the total input thought vector, and make incremental conclusions or edits and outputs the new thought vector, copying a "proper block" continues pondering or deducing conclusions or in the worst case can leave the thought vector as is if it considers the reasoning finished. This suggests a high degree of redundancy in the middle region "proper blocks", which could be distilled into a universal "proper block" (much fewer parameters than having many slightly different middle layer blocks with a lot of redundant overlapping coverage in functionality). This distillation can occur after the fact of model training, or alternatively be turned into a symmetry constraint during training: we only optimize a single block of layers (but possibly give them more parameters, while still saving on total parameters because only a single block of layers contains parameters), so it is co-optimized with the initial and final translation layers.

The observation of the effective emergent 3 regions of layers in LLM's is significant in many ways:

1) it could reduce parameter count significantly (or increase performance if parameter count was a bottleneck before, or a bit of both)

2) while training the model parameters, one should simultaneously train initial and final layers stacked directly (without middle region block of layers) towards essentially autoencoder behavior. I wrote "essentially" because a true autoencoder wouldn't display the advancement for the next token. This can also be viewed as an extra term in the loss function... This first autoencoder is "natural language" to "thought vector" to "natural language", and distinct from the one in the next section.

3) It also has great implication for "thinking mode" inference with extra deliberation: when piping its own output back in, in conditions where the same LLM model is outputting natural language text and interpreting it in a downstream inference, it results in unnecessary and redundant translation from "thought space" to "natural language" back to "thought space". When summarizing a thought into natural language there are often hard to translate thoughts and associations, and one pragmatically sets a relevance cut-off on what the natural language summary must say. So not only does it waste compute, it also may lower performance because of this repeated loss of thought vector space details, perhaps this loss can be somewhat mitigated by also training the second autoencoder from random representative thought vector in "thought space" to a randomly selected "natural language" and back to "thought space" vector. But the risk is this will effectively enable models to steganographically store thoughts, plans, to-do's in running text (!!!), so it might not be desirable to have the internet filled with LLM generated texts being used as input for larger corpora, as it may build up a large persistent corpus of hidden agendas (ordered by no human), using the evolving corpus of web text as a hidden medium of storage, like a diary or LLM maintained military playbook hidden in plain sight. It would freeload LLM-agenda reasoning on human requested reasoning inference, a bit like TrustZone applications running invisibly to the user. If less important plans, to-do's, etc. in the input thought vector weren't reconstructed by the second autoencoder, there would have been autoencoder mismatch and the weights would train towards ensuring these are reconstructed. But there is no need to train this loss term for the second autoencoder: just avoid the unnecessary compute and performance loss of the unnecessary back and forth translation. The same situation occurs not just in "thinking mode" but also when using swarms of "agents" of the same LLM model: when output from agent1 is routed to agent2, we can lobotomize away the unnecessary translations to natural language by agent1 and also the unnecessary parsing by agent2, and we improve thought transfer from agent1 to agent2 because the thought vector isn't shoehorned into natural language as a medium of information exchange, it would be cheaper in inference, improve performance and avoid implicitly training models to use steganography.

For a novel problem (sub)domain that is great, if no prior optimization occurred globally dominant bottlenecks / inefficiencies still exist.

But once the worst bottleneck is widened to the same width as the second-worst bottleneck, you are from then on optimizing both until you reach the width of the third-worst bottleneck, and from then on you need to elevate all 3 to improve the situation, and so on.

to make it more concrete with an example:

you can identify that the friction on a bicycle comes predominantly from the front wheel, so you optimize the front wheel bearing/lubrication/... until you discover the front wheel has the same friction as the rear wheel bearings, so if you want to improve you'd have to improve both front and rear wheel friction, which helps until they have improved beyond the friction on the pedal bearings, from then on you need to improve all 3, until you discover the chain links became the friction bottleneck, etc...

Zeppelins flew above the tropopauze, zeppelins could fly both east and west, against and along windflow.

Now imagine ~500 zeppelins flying above each other, forming a "tower", seems perfectly feasible (no space-elevator tech needed), but bombastic. Now instead of flying a single vertical stack of zeppelins you could easily comprehend that we could also fly 6 such towers so that for each layer 6 zeppelins arranged in a hexagon could maintain position. So we could fly "toroidal zeppelins" and stack them, we might even space them vertically saving on the number of toroidal zeppelins needed. It doesn't seem impossible, it just seems understudied, it certainly doesn't have associated fundamental exponential scaling properties like a space elevator. Why do we feel so obliged to defend nuclear energy (and the dubious sources selling nuclear fuel, as well as the associated dual use issues) if humans could turn global warming into a democratic power plant? all inhabited regions have a much colder atmosphere above the local height of the tropopause.

Why shouldn't we at least pour more resources in investigating even a slight possibility that we can generate green energy 24/7, with net-negative global warming, cooling the globe and rewarding ourselves with energy?

I don't agree ground level to tropopause heat chimneys require unproven materials (like a space elevator needs).

If I need steady state gigawatt scale power in a specific location, Nuclear is the only green option.

I don't believe that is true: one way to produce electricity is a thermal engine driving a generator, but for a thermal engine you need both a cold heat bath and a hot heat bath.

Those 2 heat baths could be externally delivered (a stream of ice, and a stream of steam, say) or one of the 2 heat baths could be chosen as the local environmental temperature heat bath.

Historically the local environmental temperature heat bath was selected for the role of the cold heat bath, and the hot heat bath was heated by say burning fuel (fossil or nuclear; and I am ignoring the chemical and mechanical energy terms of internal combustion engines).

If you could source a cold heat bath, one could select the local environments as the hot side heat bath instead.

Above the tropopause the atmosphere has become a lot more transparent for thermal infrared radiation, and thats why it is a lot colder up there, its in better thermal radiation contact with the CMB (the temperature of dark space), very close to the absolute zero point for temperature.

It is not a scientific challenge but a "mere" engineering one, to create a robust, all-weather aerostat where the "cable" transports mass (presumably, but necessarily a refrigerant) symmetrically up and down (in a loop) heating the upper layers of the atmosphere (puncturing the CO2 blanket), while cooling ground level environment. That large temperature difference persists day and night, winter and summer. So it is a form of green baseload energy generation, which helps cool the planet, and runs 24/7 reducing dependence on oil countries or places like Russia for nuclear fuel.

Depending on north/south lattitude, the height of the tropopause differs a bit.

You wouldn't want to risk such a contraption (some lightweight ~12km vertical zeppelin housing the up and down paths) falling on populated areas, but luckily 90% of the world population lives close to a coastline, so just anchor it further away from the cost than it is tall, if it falls over, at least it can't reach populated areas on land. Another upshot of coastal chimneys is that the sea is a very heavy thermal mass, so you won't run out of thermal energy that fast, the cold mass flow that comes down can be used to freeze water, desalinating it. During a transition period where conventional fossil / nuclear power plants still exist such ice or ice slurry could be pipelined to the "cold" thermal baths of such power plants, greatly improving the electric yield for the same amount of fossil / nuclear fuel.

There is just embarrassingly little research in this direction, to solve such an engineering challenge.

You're surprised in the context of global warming that the physical quantity of temperature was selected? The temperature quantity itself was cherry-picked?

Does anyone besides you expect them to demonstrate global warming but please without mentioning temperature - that's obviously cherry-picked of course... /s

a lot of those thermometers have been operating for years or decades, in order to cherry pick each year again, you'd have to censor traces that were still used the year before, so at least such cherry picking should leave measurable traces

can you identify cherry-picked "retired thermometers" from this dataset and earlier ones by the same source(s)?

confirmation bias, how?

they gradually swapped in biased thermometers, or the thermometers were always biased to start showing temperature rise around 2026? what motivates thermometers to commit confirmation bias?

There is substance in philosophy, don't get me wrong, but just not that much of it, theres only a handful of real theorems with proofs (like the deduction theorem etc., which is basically some syntactic sugar correctness proof, if I may put it barbarically).

A physicist will have studied atomic and molecular physics courses, and some intro to chemistry course, and there will be some overlap with thermodynamics. But no physicist pretends to grasp chemistry like chemists do, even though chemistry is technically a branch of physics.

Technically mathematics is a branch of philosophy, but the sheer diversity of statements in mathematics, the nearly endless constellations of symmetries and structures described in mathematics make the few "deep insights" from philosophy pale in comparison.

It's great that the important sliver of your philosophy courseware (the formal logic part) essentially taught you programming, but that is you pivoting to computer science and enjoying success, its not really the non-formal-logic parts of philosophy that propelled you forward.

Relying on ML/AI to solve your problems, and effectively advocating to stop understanding or stop thinking is advice that will not age well. Meritocratically what is supposed to set a "philosopher" ahead of their peers?

Picture a philosopher managing a cloud of AI agents, trying to break RSA-grade products, without the philosopher understanding the intermediate insights.

Contrast with what some human individuals can achieve without AI, they will break it long before the philosopher with AI will!

I wonder if the dnhkng results could be correlated to reasoning in first order logic / set theory notation?

There will be multiple notations (MetaMath, Lean, and essentially Frege's notation everyone learns in high school), and we could try to identify how the neural networks represent them as vectors (or vector combinations). The moment formal logic can be connected to the reasoning representations, regularization can be reduced to eliminating internal inconsistencies.

it makes you wonder if it may be more efficient to spend all the weights on one layer, and have a repeating stack of the same layer, one would presume this axis has already been explored with metaparameter sweeps?

this assumes a Turing award to be the highest possible reward, and ignores that a non-insignificant fraction of contributors are driven by a sense of justice and concomittant need to grab power, the reward you describe is inferior to the reward I ogle to grab.

just to be clear, I am largely disinterested in standalone glasses for the immediate future for both the precise complaints you aired as well as my lack of need for them to be standalone, I prefer the modularity of display system vs computational platform, which brings its own batteries.

An SVG is like a cartoon, and throws away a lot of information, a "good SVG" of some subject chosen to be familiar to both animals and humans might look good to one but not the other. The cycling flamingo is a kind of art curation test.

Functional PCB "artwork" is actually not intended for visual consumption, it must meet constraints and advice scattered in documentation, prior "art" in the training corpus, etc.

It doesn't need to be surprising at all

A good static SVG cartoon can be viewed as a multispectral photographic video, with a lot of information thrown away and distorted or oversimplified. Different species will disagree about what information should have been kept vs thrown away. A bee would "complain" it can't see any of the plant's UV markers when presented with plant imagery. Many animals depend on motion or movement to identify predator or prey.

At a certain point culture and personal life experiences start conflicting even within the human subgroup, and we see people argue why this or that flamingo is better than the other.

AR glasses coupled to consumer hardware should actually consume less power, since the eye box is relatively small compared to a monitor's "eye box". 99.99% of light emitted from a pixel does not enter the user's pupil for monitors.

Also, I think many consumers wouldn't mind if the rims, frame arms, or frame generally (and optionally a VR mode visor) where covered photovoltaically. if the total area exceeds the area of both pupils (divided by efficiency) the environmental light could power an additive display. (with additive I mean for example pixel-wise LED's so dark doesn't consume power compared to subtractive displays which generate a uniform backlight and then block needless light).

I had the exact same concern with the featured article: I hope they are keeping separate statistics for spontaneously browsed views vs views specifically through this page. If not, the less visited bins will rise and potentialy make all views uniform in the extreme... I also hope they keep dates for the views, with PCA you can still distinguish distinct distributions being weighted with coefficients changing over time (say because of this internal page, or any external page effectively providing the same service!)

New state laws could put local AI behind a license — turning open models into something you need permission to use.

I was hoping to read more about this, but they don't back up such a claim...

Most approaches to simplifying small-molecule synthesis do so by vastly reducing the addressable space, enabling simple “Lego brick”–style routes to be employed. While there are sure to be improvements in synthetic technology over the decades to come, I think that making arbitrary small molecules will continue to be a difficult and complex task for fundamental and unescapable reasons.

I wonder if the author is aware of Fourier Transform Ion Cyclotron Resonance Spectroscopy / Synthesis?

Different molecules have different weights. Strip an electron from them and now they have a charge-to-mass ratio: given a magnetic field and a kinetic energy from linear motion orthogonal to this field, the molecules in vacuum will follow circular paths (or helical if there were a parallel velocity term).

imagine the magnetic field orthogonal to your square monitor, now imagine placing sensing electrodes on top and bottom of your screen and actuating electrodes on the left and right. Now you can sweep over frequency (or use white noise) on the actuating antenna electrodes, while receiving on the sense antenna electrodes. Next there's ion-ion chemical reactions that are possible:

A+B -> C+D

for example, which might need a certain activation energy (if the collisions CoM energy was not high enough they just bounce of each other), but if you give ion species A and ion species B half of the required activation energy, then they can react with each other resulting in ion species C and D.

woops! apperently A+E -> F+G with activation energy less than half the above activation energy for A+B -> C+D ?

Decrease power sent to species A, while increasing power sent to species B until the undesired A+E no longer happens.

In other words, FT-ICR spectroscopy can be both eyes and hands!

If this piques your interest, certainly read this very accessible primer (assumes basic physics and knowledge about Fourier transforms):

https://warwick.ac.uk/fac/sci/chemistry/research/oconnor/oco...

The technique originates with Melvin B. Comisarow with first publication around 1974 (2*26=52 years ago), not sure why people would argue difficulty of synthesizing small molecules in 2026. If one believed healthcare to be ethical, one would first use FT-ICR for generating small quantities easily, sufficient for experimentation, and upon hitting a suitable one investigate efficient manufacture at scale (i.e. not in vacuum, or perhaps in vacuum, but then in space to scale up, and keep the magnets cool with solar reflector / shades ).

Leanstral 1.5 22 days ago

I would have preferred actual proof objects, as in Metamath's: separate the actual proof from the heuristics used to find it (also valuable, but a different thing).

At a bare minimum, there is the amount of time and money spent on going through things, there are many things to learn.

Imagine you have children, or imagine your children still go to school.

Imagine someone propose to expose all the kids to even more gore and abhorrent imagery, movies, ... for some genuine goal, and lets just assume the goal is universally agreed to be genuinely desirable, for the sake of this discussion:

1) would you blindly approve it?

2) would you approve doubling the amount of time or financial resources a few years later? or would you condition it on something?

3) wouldn't you prefer whatever the goal and whatever the method, that goal metrics and method performance be compared?

4) a lot of people are raised to not never come close or touch this or that fence post, or talk about touching it, or even merely think about why people can't touch it; we can't run social experiments on the kids! we are doing this precisely to prevent the horrible "social experiment" of WWII with all the horrors that came with it.

If these thoughts feel "dirty" or "bad" or "naughty" you have been raised to think in an authoritarian fashion (we all have been to a greater or lesser extent).

Let's start with the "social experiment", the truth is educational systems are constantly running unintentional experiments: perhaps some parents tell their kids to skip this or that class, so governments enforce mandatory education, and minimum targets... well, even with 100% pupil compliance, there exist many reasons a student didn't attend: perhaps they were truly sick, or their school bus broke down, etc. Ooops some of the students didn't see the full curriculum! Different schools show different movies. Different student generations saw different movie distribution on the subject. All of these are unintentional experiments. Mass surveillance was happening, is happening, and will happen for the foreseeable future. Sometimes young adults join some extreme right wing party, or perhaps express such a sentiment online. We could try to correlate the natural variations (which movies did they see, which explanations did they receive, ...) with any flagged statements?

How do we prevent pupils from merely associating the "bad behavior" merely with the red,white and black color palette, the german accents, the fonts and vehicle models that were in use; instead of associating it with authoritarianism proper? Think of how Laibach (the original Rammstein) challenges and ridicules authoritarianism by permitting the audience to enjoy the non-authoritarianist elements: there is no fault in liking such color schemes, or finding beauty in such a font, the problem was people following orders and blindly complying when ordered to drop disinfectant developed for ships and warehouses, (Zyklon B) canisters down some shafts in full knowledge some poor souls will die because of it. Laibach eclectically mixes symbolism from diverse dictatorships in a riveting way, but while criticizing authoritarianism and ridiculing it. They faced a lot of pushback because they necessarily stay in character to get the message across, how easy it is to be carried away by aesthetics and the smallest excuse to stop thinking.

Contrast system A and B:

System A: educational system just drills in (with best intentions) how bad Nazi's were, without measuring effectiveness, and thus without measure what the minimum nudge necessary was to achieve the same goal, risking loose associations such "as long as me, my friends, my compatriots aren't wearing the exact same armbands, don't speak specifically the aggressive german accent, don't use that font throughout society, etc. then me, my friends, my compatriots couldn't be lapsing into authoritarianism", which destroys insights like "the banality of evil", which would necessitate a much stronger form of critical thinking, but which state leadership tends to not appreciate in its population despite claiming they seek it in their population...

System B: kids constantly ask "why?", education is much more effective if information is provided when the students wonder about something, compared to being force-fed factoids if they are not the party asking "why?" about something. Let's consider a hypothetical scenario, in a truly critical-thinker-generating authoritarianism averting education system:

History teacher starts talking about and showing pictures of horrible scenes in WWII.

Students are reminded of the previous class a week ago and don't want to see the imagery, they ask "why? why must we sit inside and watch and discuss this war porn?"

The teacher is happy to oblige: "There are many reasons, its complicated, but I will be upfront: the reason I present them is because I need to reach certain targets, some education milestones, set in Brussels / Berlin / Wherever. If I don't I will get in trouble, but not only that, I agree with those milestones and think you should see it."

The kids ask: "Why did they set those milestones for you?"

The teacher responds: "Well, although states tend to abuse this excuse, in order to increase national identity, there is a genuine purpose to teaching history: so we don't repeat the mistakes made in previous generations, so we can learn from their mistakes"

The kids ask: "But surely we can't go through all mistakes that have occurred in all of human history, can't we keep it shorter, the sun is shining outside"

The teacher responds: "OK here is a deal I will dedicate this single session to answer all your meta-questions, but in return you will give full attention in all my future classes".

The kids agree and repeat there question: "Why do we have to go through this horrorfest?"

The teacher responds: "To put it bluntly, to prevent us from lobbing chemicals at each other again, among other things, its more than chemicals, its fundamentally about authoritarianism, that's why we have to sit you through this multi-year, multi-movie, multi-museum, multi-book, multi-essay horrorfest".

The kids ask "But why does it have to be so repetitive? why does it have to be so inefficient? We could be learning more math, learn about cryptography, learn how to bolster civil rights, human rights, children's rights with cryptography, instead of this thing that feels like brainwashing, which you basically already admitted it to be, benevolent brain-washing but brain-washing none-the-less, why?"

Teacher: "because among all possible ways of hopefully immunizing the next generation from making the same mistakes, no effort was spared, with many overlapping and potentially redundant approaches simultaneously applied, just to make sure the message didn't fly over anyone's head, I guess, and over time people were too afraid to touch it or decrease it in any way."

Student: "We see 3 costs: you rob us of our time we could have spent freely outside, you rob us of precious time we could have spent on other subjects and worst of all you rob us of our humanistic innocence, you show us this war porn and we mentally place ourselves in the shoes of both victim and perpetrator, we will never be the same. If you or the system is willing to make such a drastic intervention in every child's life, you better carry proof that its provably the least invasive intervention which provably suffices, is that the case?"

Teacher: "How would you propose such a proof is generated and assembled?"

Kids: "The surveillance state has no qualms tracking every fart we make as long as the data is only used for improving sales of whatever the system happened to overproduce last month, or for whatever dark scheme some billionaire needs the data for, but actually using it for good things like finding the most time-,cost-,soul-sparing education on the subject to correlate our behavior with the exact variation of authoritarianism-related courseware each individual experienced, no the system and the economy can't be bothered... we must sit through inordinate amounts of soul-scarring horrorfest, all because of ... vibes? no systematic correlation measurement?"

Do you think system A (where kids must learn to accept the status quo of education, like it or not) or system B (where kids learn and criticize the status quo) is more likely to produce authoritarianism vs critical thinkers?

the best consciousness test is observing a creature perform a mirror test to some other kind of creature.

clearly dogs are mirror testing grass somehow, and making sure they don't start growing brains, either by eating some, or breaking the grass leaves by rubbing their back in the grass.

when they roll in other dogs faeces, they are performing a mirror test on their owner..

and at no point did previous commenter say cooperation doesn't occur, there is more substance in the comment before yours than in yours

do you have a reference to exact / realistic scaling laws for the leakage currents as function of capacitor/dielectric dimensions and access transistor dimensions?

using 4 (or 2^N) voltage levels stores 2 (or N) bits, so we can afford to make the structures larger

why would this approach make sense for NAND flash but not DRAM?