HN user

rsfern

1,659 karma

Working as a researcher in materials science (physical metallurgy specifically).

My research interests include applied machine learning for quantifying materials substructure, which we call microstructure.

Posts17
Comments705
View on HN

My point with the force field example wasn’t to argue against neural scaling as a valid strategy, it totally is effective and a lot of groups are doing it. But I feel like we might be talking past each other a bit.

What I’m pushing back on is what I think is a sort of one-dimensional view of Sutton’s bitter lesson. People seem to equate it with model scaling, but there are lots of general ways to leverage computation that don’t involve just scaling models and supervised training datasets up. For example Sutton’s first example is straight up search, no parameters at all.

The point of the force field example is that it seems you don’t need billions of parameters to represent the functions we’re interested in, but with small models it’s harder to find those functions by pushing harder on the standard training algorithms, and that maybe some different algorithm that leverages computation more effectively could do so.

I don’t think there’s a fundamental reason that performance has to be monotonic in model size or even training FLOPs. At least I don’t think it’s been proved to be so, so I think “misinformed” is a bit premature and sort of makes GP’s point.

There’s evidence that model size and representational capacity are not exactly the same, and that scale is maybe more important for learning than it is for representation (past a point). Consider the early work from the current neural scaling paradigm. The Chinchilla scaling study shows that smaller models can match the performance of larger models by training longer.

To GP’s point, if everyone is exploiting the scaling lever, few resources are being allocated to finding more efficient training algorithms that could let us work with right-sized models instead of pulling the scaling lever as hard as we can afford to.

I’ll end with a dramatic example from my field of materials science (which admittedly might not strictly generalize to LLMs). A lot of the field is pursuing the model scaling strategy, and it’s still paying off. But [0] recently reported competitive accuracy with much smaller models that run faster and can address much larger problems. The model architecture is pretty much the same, but they use a different training strategy and really focus on data quality

0: https://arxiv.org/abs/2504.21286

I don’t mean to pick on you in particular here, but this approach has been bothering me a lot lately, and it seems like it’s super common.

I get that it’s an early prototype and not all the design choices are made yet, but I struggle with “I can’t afford human-readable documentation yet”. Isn’t human readable documentation important for efficiently planning and deciding what you want to build? I feel like I can’t afford not to have human readable docs and plans while the project is taking shape

Related, if a user can prompt an agent to translate to a human-readable summary, wouldn’t it be better to just do this in place? Sure, models can deal with noisy LLM outputs but shouldn’t a document that’s easier for humans also be easier for bots?

I often start prompts with “please”, but I usually don’t thank the model. Framing a question or a request for help with “please” is in distribution for me, it’s a distraction from composing a thoughtful prompt about my actual question to go back and edit out politeness.

I don’t reply “thanks” like I would to a person though, I just close the chat if I have no more follow-ups

The DHS secretary seems to me to have the point of hosting international students backwards

This final rule ensures that foreign students remain focused on their primary purpose: completing their studies and returning home.”

Especially at the PhD level, why would educating foreign students and sending them away be the primary motivation for granting visas? Historically it’s been about attracting the best and brightest in the world, training them to be excellent researchers and scholars, and giving them a path to becoming citizens and contributing to growing our economy and technology development and all that. Sending them away after investing in training them for years makes no sense! Especially given the administrations adversarial stance on technology development as a competition between nations more than as an opportunity for collaboration.

I mentor a postdoc who might be affected by this, he’s on a J visa and it’s already been a nightmare of paperwork. He’s really good. Not disputing that there’s some amount of people gaming the student visa system, but this just doesn’t seem well thought out to me :(

True. But it has no idea that it has no idea, so it might be able to look back at the session trace and pattern match its way to actionable feedback?

I think JEPA is super interesting, but I feel like this example highlights some of the challenges of long horizon planning. For one, chunking the planning stage into a bunch of intermediate goals seems really limiting, because a lot of what makes model based control interesting is that we don’t want to impose a solution strategy (because we want to solve problems we don’t know how to solve)

Another thing that has been bothering me is that you have to write the goal in input space. That doesn’t align with all problems, for some problems there could be many different states that satisfy a goal. For Mario maybe it’s ok, but there’s some weirdness still, like should the goal state be Mario at the finish line of the level with a specific timer state in the frame header? What about optimizing the number of points?

Also it’s interesting to think about how you would get Mario to reliably jump on koopas and goombas. IIUC JEPA models are usually trained with random rollouts, and then you’d handle this sort of intermediate goal in the planning optimizer? But that seems inefficient, and including some planning in the pretraining rollouts might be necessary to get enough relevant intermediate states. And then it starts feeling like reinforcement learning…

I’d be happy to have a check on my intuition here, or pointers to interesting writing on these topics

p.s. on topic, I liked the debugging strategies used in the blog post, that was my favorite part of the writeup

What aspect do you consider basic? I haven’t had a chance to read more than the abstract because of the paywall, but the really interesting thing here is the mechanically induced transformation that leads to a 3-phase nanocrystalline alloy. I haven’t understood the “fully coherent” part from just the abstract, but I think it’s a very novel report. It means they have three different crystal structures with seamless interfaces because all the atoms line up at where the crystals meet. To do that with three structures is remarkable. I’m not sure if it’s the first example, but I’m only aware of alloys with two coherent phases (some superalloys for example)

There’s some precedent for mechanical deformation to get nanocrystalline grain structure, and some precedent for mechanically induced phase transformations (see TrIP steels) but I consider both of those concepts pretty advanced

TrIP is probably the closest thing, but I’m not sure how widely known there are among the “metal forging” community? TrIP is usually targeting well known phase transformations to two-phase microstructures, here we have three nanostructured phases.

Finally high entropy alloys are absolutely not well understood, even if the idea of mixing a lot of elements and getting a disordered solution seems simple on its face.

TrIP: https://en.wikipedia.org/wiki/TRIP_steel

This is really cool metallurgy. They start with an alloy and deform it and because of elemental size mismatch they can cause the alloy to self assemble into nanoscale crystals with three different structures

The paper: https://www.science.org/doi/10.1126/science.aec4995

As an aside, “super alloy” is not the best wording choice on the part of the author of this sciencealert article, superalloys are an established alloy family that follow a different design strategy and have a very different composition profile https://en.wikipedia.org/wiki/Superalloy

But this is true for lots of solids, not just glass. Consider creep deformation [0] which is deformation mediated by diffusion of defects in solid materials. It’s a big problem in metal turbine blades, it limits the maximum usable temperature to well below the melting point.

The physical mechanism would be similar in glass flowing in this way, so I don’t think evidence of glass flowing like this should make us think of it as a liquid instead of an amorphous solid

0: https://en.wikipedia.org/wiki/Creep_(deformation)

Lab to product scale-up is a well known hard problem in materials (and chemistry), lots of public and private investment has been aimed at accelerating this for decades

I think the distinction between discovering a material and processing / scaling it up is a bit artificial. A lot of people think of a new material as just the crystal structure or something, but really all the defects and complex multiscale structure is just as much part of what defines a material, and controlling all that is why materials development is hard, and why you need so many different complementary measurements to understand what’s going on

I was a bit underwhelmed by this writeup because it’s a bit generic. I didn’t really see any specific new ideas on how to accelerate this process, or to differentiate from the main stream of materials discovery research which has been pretty dominantly AI forward for at least five years now

EDIT: I checked out some of their case studies and they’re pretty interesting and exploring some new materials characterizations territory. They’d be more impactful if they were more than just text IMO but much more concrete and less generic than the linked post

Maybe I have too optimistic a mindset, but “just be honest” in academia isn’t about being a rule-follower, it’s about not short-changing yourself by coasting though on autopilot instead of learning to think and solve problems for yourself.

Whether that really matters if your goal is to climb the social ladder and have power and influence, I don’t know.

You allowed a take-home exam which means students are able to use any and all resources.

It was a closed-book exam. The professor shouldn’t have to hold students’ hands for them to act with integrity, they are all adults.

In this particular class, the professor made the final exam in-person, and didn’t count the take-home midterm because the score distribution wasn’t consistent between the two exams. I think that’s a reasonable approach, but it’s kind of sad that it was necessary

If the explicit role-playing prompt is just to identify multi-valent terms, then revising the question to include more specific context without a role-play prompt should work just as well right? I’d be really interested if anyone has evaluated that hypothesis

A fun (frustrating) feature of language is that we get these name collisions even with a single domain. One that I have to remember to revise myself fairly often these days when chatting with other experts in my field is “diffusion model” which can either mean generative deep learning or a differential equation describing mass transport.

High-Entropy Alloy 2 months ago

That’s the active research area GP mentioned. In startup land there are a few large outfits, Lila Sciences, Period Labs, Radical AI are all doing a mix of simulations, AI, and autonomous laboratory infrastructure specifically for materials science. (Lila does a lot of biotech but the have materials researchers too)

Also lots of interest and activity in this space in the national labs and academic research scene

High-Entropy Alloy 2 months ago

It depends what you mean by commercially interesting. There’s loads of interest in aerospace (for high temp corrosion resistant structural components) and catalysis but these alloys are pretty much across the board at a relatively low level of technical readiness. It’s developed enough that there’s significant industry R&D and not just academic and government research, but I don’t think there’s really wide-scale deployment yet of alloys with 4+ principal elements

It seems like an opportunity for a hierarchical cache. Instead of just nuking all context on eviction, couldn’t there be an L2 cache with a longer eviction time so task switching for an hour doesn’t require a full session replay?

I think the exceptionalism is the other way around. What makes anyone think they understand what makes for intelligence when we barely understand our own neurology?

Admittedly my understanding of QM is a bit vibey but I’ll try to answer

In an atom, angular wavefunctions with wavelengths non-integer divisions of 2pi can’t exist because of the boundary conditions on the wave equation. A free electron can have any wavelength, but once you put it in a box (confine it to the potential around a proton in a Hydrogen atom) the non-integer wavelengths aren’t allowed

I think it’s instructive to think about what the wavefunction represents. It’s square is the electron probability density (technically the wavefunction is complex valued so it’s the wavefunction times it’s complex conjugate). If you have a non-integer multiple wavelength then the wavefunction goes out of phase with its complex conjugate after one period, and if you integrate over the angular domain the electron probability has to be zero everywhere.

This also answers your second question. The radial solution to the wave equation for hydrogen gives you the Laguerre polynomials. They don’t all go to zero at the nucleus though, actually the first one has a maximum at zero because it scales like exp(-r) (See fig 4.10.2 on chem.libretexts linked below). But when you do a volume integral to calculate the electron probability, the probability near the nucleus is low because the integration volume is small even though the wavefunction is large

https://en.wikipedia.org/wiki/Laguerre_polynomials

https://chem.libretexts.org/Courses/University_of_California...

The reason is that electrons (like all quantum mechanical objects) are wavelike. In an isolated hydrogen atom, the electron is in a spherically symmetric environment, so the solutions to the wave equation have to be spherical standing waves, which are the spherical harmonics. The wave frequencies have to be integer divisions of 2pi or else they would destructively interfere. (Technically each solution is a product of a spherical harmonic function and a radial function that describes how fast the electron wave decays vs distance from the nucleus)

What’s interesting is if the environment is not spherically symmetric (consider an electron in a molecule) the solutions to the wave equation (the electronic wave functions) are no longer spherical harmonics, even though we like to approximate them with combinations of spherical harmonic basis functions centered on each nucleus. It’s kind of like standing waves on a circular drum head (hydrogen atom) vs standing waves on an irregular shaped drum head

Of course the nucleus also has a wave nature and in reality this interacts with the electrons, but in chemistry and materials we mostly ignore this and approximate the nucleus like a static point charge from the elctrons perspective because the electrons are so much lighter and faster

While the treatment for methanol poisoning indeed includes ethanol, I don’t think your dosage suggestion is right. Your body would still have to process all the methanol, the job of the ethanol is just to slow down the reaction. If you suspect methanol poisoning you need the hospital, they will administer the ethanol intravenously and I think do dialysis to remove the methanol and the formic acid it metabolizes to (this is one of the toxins in ant venom)

https://doi.org/10.1053/j.ajkd.2016.02.058

There are groups that are actively working on automating conventional labs like this. Most of the efforts I know about use non-humanoid mobile robots or even just a six-axis arm on a rail and some lab space reconfiguration

This issue of accessibility is widely acknowledged in the academic literature, but it doesn’t mean that only large companies are doing good research.

Personally I think this resource mismatch can help drive creative choice of research problems that don’t require massive resources. To misquote Feynman, there’s plenty of room at the bottom

I like this analogy of always choosing “I’m feeling lucky” on Google, I feel like it clarifies a boundary between information retrieval and evaluation that gets blurred by language model summarizations. I’ve been frustrated with the LLM summary at the top of the Google search results for scientific topics because often the sources linked to don’t actually contain the information the summary is citing them for. Then I have a side quest of finding the right backing literature or deciding the summary was just wrong in the first place

I don’t know why the author of the article wrote “could”, but I personally work closely with some non-high-risk-country NIST foreign guest researchers. It’s been filtered down verbally through the management chain that the end of this September is the re-review deadline, and it’s not been stated as a hypothetical.