HN user

btrettel

4,190 karma

I am a mechanical engineer working in computational fluid dynamics.

Personal website (including contact information): http://trettel.us/

Posts0
Comments1,170
View on HN
No posts found.

Some comments not addressing your question:

I think you are holding climate science to a far higher standard than the other physical sciences.

In the physical sciences, blinded experiments are rarely done. The need for blinding is greatly reduced compared against the medical and social sciences as the data measured is far more objective and observers usually can't influence the results. I've done experiments where I basically start recording data, turn a valve, and from that point on, the result is outside of my control as long as I'm watching from a distance of 10 feet or so. That's often not the case in the medical and social sciences. In my experience, blind experiments in the physical sciences take the form of a blind prediction challenge where a bunch of teams are asked to predict what a certain experiment will do before the experiment is run. This is a good practice that I advocate. I also am interested in blind data analysis and "fake-data simulation" as Andrew Gelman calls it.

There also are a lot of times in the physical sciences where duplicating the "full scale" case is not possible for various reasons (cost, legality for nuclear weapons testing, etc.). Instead they do experiments on scale models (which might not have similitude [0]) or parts of the full problem. This is not ideal, and I do think the researchers could do better, but the situation is unavoidable. Again, there's no reason to single out climate science on this.

I also want to strongly push back on the "computers models are basically mathematical assumptions, not physical laws" part. I'm a mechanical engineer who works in computational fluid dynamics. The models I use have a lot of overlap with climate models, but don't get as much scrutiny even when they are of similar reliability. A large computer model like a climate model has a lot of components, some of which could be regarded as very reliable (likely what you mean as "physical laws") and some of which are less reliable. But even the less reliable components are not assumptions. They always are backed up by some data, perhaps not as a comprehensive as is wanted, yes, but some data.

[0] https://en.wikipedia.org/wiki/Similitude

Drag-and-drop of the PDF file worked for me in Zotero 9 (latest version). I never used this feature before. This would be greatly preferred to my earlier suggestion to get a LLM to generate a list of DOIs.

Zotero is not just a front-end to BibTeX. Here are some things I like about Zotero off the top of my head:

Zotero integrates well with various online services. This I think is Zotero's most valuable feature. The simple fact is that not everything exports BibTeX. Can you get a library catalog to export BibTeX? Perhaps in CS, most publications provide good BibTeX, but a lot of journals I (a mechanical engineer) deal with don't provide BibTeX at all to my knowledge. But I can simply press a button in Firefox and import the bibliographic data embedded in the web page into Zotero. (No, Google Scholar is not a good solution here because it's frequently inaccurate and incomplete.)

I'm a technical guy who can handle BibTeX fine, but I still prefer the Zotero UI over using BibTeX files in a text editor and shell, even if I'm only generating BibTeX files.

The author discusses using non-standard fields like keywords to store extra data. I would recommend that to store additional context about the document. Zotero can store even more than that, including web pages, files, and additional notes. I like that Zotero saves a snapshot of the journal article web page in most instances. Journals sometimes do go offline, so it can be nice to have the web page. I often have detailed notes about particular documents in notes in Zotero. Yes, you can do that with BibTeX, but I could see the extra fields cluttering the BibTeX file. (Contrary to what the author states, the note field probably should not be used as they describe because it's printed in some/most? bibliography styles.)

The search in Zotero is more powerful than grep. Try returning bibliographic entries where one field contains X and another field contains Y. I think you can probably come up with a grep solution for that, but it's so convoluted that it's probably never used in practice. You could use a BibTeX searching program like biblook to get around this problem.

Don't get me wrong. Zotero is far from perfect. It can be slow, and at this point I would prefer a TUI reference manager. But overall, it's the best option I've tried.

First I used Claude Code to generate all the BibTex files for the PDFs

I think this has an unnecessary risk of hallucinated bibliographic data. For anyone doing something similar in the future, it would be more reliable to make a LLM generate a list of DOIs and have Zotero import the DOIs.

I would suggest looking at more powerful file managers. I use Midnight Commander, which isn't perfect, but is far more powerful than what is provided by default and is also extensible through things like its user menu. My organization system makes heavy use of various Bash/Python scripts and Midnight Commander user menu items to make the file system work for me.

***

A friend of mine has a similar single folder PDF organization system and lately has been trying to better organize it. As I understand it, he's keeping the one folder, but using a reference manager software to track metadata like tags. In contrast, for about 15 years now, I've stored various documents (mostly PDF files) in a folder hierarchy with Zotero for bibliographic data. The PDF filenames match the Zotero citation keys so I can easily switch between the two systems. I'm currently at about 39K PDF files in 4.5K folders with 17K symlinks.

My friend has slowly chugged along, organizing perhaps hundreds of files manually per month. (Maybe he's stopped, I don't know.) It seems to me that transitioning from a one folder approach to something with more context (whether metadata or a folder hierarchy or whatever) is really hard. The best time to add that context is when the file is obtained as I might not remember later. I personally don't trust a LLM to spot the nuances that matter to me when organizing things. A generic organization scheme is of no value to me. I don't know what the goals of your organization effort would be, but this is something to consider.

I think part of the problem is that the people hiring are looking for shortcuts and won't accept any approach that requires more than a certain amount of time.

I've read comments about how X company got hundreds or even thousands of applicants for a position, but they can't possibly look through all of them. Well...

When I worked at the USPTO as a patent examiner, I never had that few documents to search. If I said that I had only 1,000 documents to search, my supervisor would probably be suspicious that my search was poor quality. I was able to find prior art within the first couple hundred documents searched at times, but that doesn't mean that the search could stop at the USPTO. Stopping early like that can come back to haunt an examiner later on when the attorney revises the application. There were many applications where I closely examined hundreds of documents, often looking at drawings and/or text for particular features with a list checked for every single document. I can think of one water heater patent application where I probably closely examined thousands of documents looking for a particular shape of a component, and eventually found it in a fairly obscure Korean patent!

You can complain about the USPTO's quality all you want, but they're at least doing a better job than HR, who simply gives up and says they can't find it when it's right there if they actually put in the effort. The USPTO rejects nearly everything with prior art in the first action.

I'm a mechanical engineer who has written similar tools for work and hobbies. Producing pretty pictures does not mean that the model is physically accurate. Unfortunately, such tools seem be evaluated much more on flashiness and not on more reliable and objective criteria like physical accuracy based on verification and validation test suites. I'm seeing that in the comments here. I don't think LLMs make what I do irrelevant, but I have thought that I'm going to have to improve how flashy my simulations look to compete better with non-experts who use LLMs.

Unfortunately, running an online forum has been a pain for a long time. You have to promote the forum to keep it active, deal with various bad actors (historically spammers, trolls, and black hats), maintain the forum, deal with drama, etc. It adds up and largely isn't appreciated. People act like a forum exists by itself, but it doesn't.

And in the last 5 years, running an online forum has become more of a pain given how badly behaved some bots are. Just recently, I installed Anubis on some online forums I run, and I've been amazed by how much traffic dropped. Before, server load was becoming a problem to the point where all the forums I ran were taken offline by their hosts! I have been thinking about how there's a need for a forum software which produces static HTML for the content, while all dynamic components are behind a login. Bots won't increase server load much in this case. If the forum administrator decides to end the forum, they can easily keep hosting the content without any future maintenance beyond paying the bills. Two of the forums I run are just archives at this point, and I'd love to be able to flip a switch and make them static HTML... (I probably will adapt some script I found on GitHub do to this in the future, maybe with help from LLMs.)

Location: United States (Open to any US location)

Remote: Yes, open to remote, hybrid, or in office

Willing to relocate: Yes

Technologies: Fortran, Python (Matplotlib, Numpy, Pandas, Scipy), OpenMP, Git/GitHub, Linux, Bash, others...

Résumé/CV: Available on request

Email: cxrqnw5z@trettel.us

GitHub: https://github.com/btrettel

Personal website: http://trettel.us/

I'm Ben Trettel, an experienced mechanical engineer with a PhD, specializing in computational fluid dynamics, design optimization, and verification & validation of computer simulations.

I am particularly interested in opportunities to build cutting-edge physical products where computational simulation and design optimization are key.

A spell checker, grammar checker, and tutor change a relatively small fraction of the writing, preserve the writer's style/voice, and rarely introduce errors that are hard to detect like hallucinations.

A translation app changes nearly 100% of the content, often changes the writer's style/voice, and can introduce hard to detect errors. But there's a far closer correspondence to what was written by the original writer. The basic ideas are still from the writer. A translation app is not expanding a short idea into something longer, and including some things the original writer never thought in the process.

***

Pre-LLMs, I did in fact disclose when I was using a translation app in some translations of scientific articles I produced. It would be weird to disclose the use of spell checking, grammar checking, or who previously taught me writing as these things are ubiquitous. I will also acknowledge people who were influential in my thinking. If a LLM is doing a lot of the thinking for me then I do think disclosing LLM use is appropriate.

Recommender systems for papers tend to be pretty bad, so there's a lot of room for improvement. I'll use Semantic Scholar as an example. I have a bunch of folders in what they call a "Library" with recommendations turned on. Semantic Scholar tends to recommend things that are in the same general area but not specific enough. So I guess that Semantic Scholar seems to interpret adding a paper to a folder as expanding the scope of the folder, but it could be narrowing. There's no way to distinguish between the two. Their recommender system is supposed to magically figure it out. Some way to add additional context like relevant keywords or a way to select which parts of the papers are relevant would be helpful. As it stands, I have to repeatedly thumbs down recommendations, and Semantic Scholar doesn't figure out what I mean from that vague signal and instead stops recommending much anything. It's not that there are no additional papers to go into these folders either as I've added more over time that I've found through other means.

Unfortunately, when managers expect a certain throughput, solutions like this often appear. I saw similar systems when I was a patent examiner at the USPTO. From what I recall being told, the USPTO's "form paragraphs" started out in the DOS era as some WordPerfect macros developed by an examiner, not management. Management defined the quota and examiners came up with creative solutions to meet the quota.

These solutions are symptoms of a broken system. I would not judge the people working within the system for using these solutions (edit: unless someone's quality is exceptionally bad like academic fraudsters as a academics work in a similar metrics-driven environment). Management (and politicians in the case of the USPTO) created the incentives that are the real problem.

Automation like this can be useful to enhance quality, but in my experience at the USPTO, there was a lot of automation they could have used to enhance quality, but they didn't. The incentives to improve speed are far stronger than the incentives to improve quality.

Getting Western researchers to not ignore papers published in languages other than English is a serious uphill battle.

During my PhD, I tried to do a truly global literature review of all languages. I tried to be a lot more comprehensive than others before. I got quite good at locating existing English translations of non-English papers [1], and even published around a dozen English translations I produced using Google Translate, DeepL, and Yandex Translate [2]. My PhD advisor clearly disliked this aspect of my research and didn't seem to consider it to be research. Despite that, I remember when one of my papers was being reviewed that a reviewer commented that X was the most valuable contribution of my paper. Problem was that X arguably wasn't a contribution because it came from a Russian paper published back in 1963, as I stated in my paper! Yes, I did a better job at X than they did back in 1963, but it's only a "contribution" because people ignored non-English papers.

Ultimately, I think ignoring non-English papers is one aspect of the larger problem that literature reviews tend to be non-comprehensive. The literature is rarely well organized, and no, LLMs don't solve this problem, though I do think they help and will get better over time in this aspect.

[1] https://academia.stackexchange.com/a/93209

[2] For example, here's one that I find was way ahead of its time (original was published in 1938) and still arguably had publishable elements back in 2020: https://repositories.lib.utexas.edu/items/ca7fc7d3-cf16-4859...

I don't think I experienced discrimination during admissions either. Off the top of my head, I don't know any US citizens who told me that they wanted to go to grad school but were unable to be admitted to a school.

As a US citizen with a PhD, I didn't experience any clear discrimination in favor of foreign students during grad school.

I think the main reason so few US citizens get PhDs is because PhD "student" (they're actually workers) positions pay so poorly. Make PhD student positions have non-poverty wages and you'll see a lot more interest from US citizens.

On the flip side, I think foreign students experienced a lot of abusive conditions that I could more easily say no to because I didn't have a visa that required me to work at the university. I've seen some of that first hand. I don't mean to imply that there would be no cost to me saying no, just that I wouldn't have to leave the country if I said no.

Even if it wasn't a large size, it likely wouldn't be great. During my PhD on sprays, I did some (unpublished) experiments using isopropyl alcohol to reduce the surface tension. The nozzles I used were around 1 mm in diameter as I recall. I did not anticipate that the room would fill up with isopropyl alcohol vapor and (probably) tiny droplets. I wore a mask and maybe left the room while each trial was running. Breathing that likely wasn't great for my lungs.

Fluid dynamicist here. The word "compressible" has multiple meanings and this might be confusing you. You don't need compressible flows in the sense of high Mach numbers. There are other models where the flow is variable density, but thermodynamic and hydrodynamic pressure are decoupled to remove the pressure waves that make high Mach number flows hard. There's also the Boussinesq approximation for buoyancy when the density varies only a small amount. I'm not particularly familiar with atmospheric models, but I'm sure they don't use the high Mach number form. "Incompressible" methods are common for the second class of model I mentioned, though how to use them so might not be obvious.

Location: United States (Open to any US location)

Remote: Yes, open to remote, hybrid, or in office

Willing to relocate: Yes

Technologies: Fortran, Python (Matplotlib, Numpy, Pandas, Scipy), OpenMP, Git/GitHub, Linux, Bash, others...

Résumé/CV: Available on request

Email: cxrqnw5z@trettel.us

GitHub: https://github.com/btrettel

Personal website: http://trettel.us/

I'm Ben Trettel, an experienced mechanical engineer with a PhD, specializing in computational fluid dynamics, design optimization, and verification & validation of computer simulations.

I am particularly interested in opportunities to build cutting-edge physical products where computational simulation and design optimization are key.

I think the type of speed that you're referring to is different from the type of speed the linked article is referring to. I'm not sure what the best way to distinguish the two is. With slow software, you're presumably getting a right answer, just slower. In my job, people who work quickly often produce a wrong answer.

One approach for testing with multiple compilers that I use on some Fortran projects (where testing against multiple compilers seems more common than in C) is to use a variable from the command line to specify the compiler, for example:

    make FC=ifx check
On my Fortran projects, that will run the tests with Intel's Fortran compiler. The Makefile has logic to automatically change compiler flags as appropriate. I default to the GNU Fortran compiler, so `FC` isn't required.

I have made a script to run through a series of compilers by alternating between `make check` and `make clean`.

I have separate Makefiles for GNU Make and NMAKE/jom. My Fortran code works fine on various Linux distributions and Windows, though I'll add that achieving that is probably easier with Fortran than C. I've also tried a BSD Make that worked (on Ubuntu at least). My Makefiles are pretty close to the intersection of POSIX and NMAKE, so the main differences between the different Make versions are the conditional statements needed to handle the different compiler flags and the include statements (as I put the compiler flags in separate files).

I had a similar setup in 2023, but the computer was reformatted after I moved. I wrote a HN comment about the setup before: https://news.ycombinator.com/item?id=37792204

I liked it and intend to use a similar setup in the future. There were quite a few "rough edges", unfortunately. In retrospect, a tiling window manager would have been a better choice.

I found Midnight Command to be great for this, with its integrated file manager, file viewer (mcview), editor (mcedit), and diff (mcdiff).

I didn't realize how much I relied on a unified clipboard until I didn't have one any longer. mcedit's clipboard was a file (or one of them was?), so I had to adjust some workflows.

The biggest problem came from my need to view a lot of PDF files. I had a framebuffer PDF viewer that was pretty clunky. It did not work with tmux and PDF files could not be opened directly from Midnight Commander as I recall. This specifically is why I'm thinking about a tiling window manager as I won't have to pick a clunky PDF viewer and the remainder will just work.

Thanks for the reply. You're right that the data for this is very fragmented. Victor was looking at Crossref metadata. I think he always had what he was doing on Codeberg, though I'm not sure. I was looking at arXiv and 1960s to 1980s printed translation indices listing translations on paper that are today in archives uncatalogued at the Library of Congress, British Library, and other libraries/archives. (The indices list which libraries have each translation and what it says is accurate for the Library of Congress in my experience.) OCR was not cooperating on turning my scans of the translation indices into something I could parse, despite the indices having a regular structure indicating that they were computer-generated. LLMs likely would help with that now, but all of this was pre-ChatGPT. My plan was to automatically convert the bibliographic data in the indices to DOIs, but as it turns out, a large fraction of the articles in the indices do not have DOIs. We ultimately did not consolidate these sources.

Anyhow, it's obviously a huge task and I don't expect you to build this. I was just curious if you had thought about it as you clearly have a lot of relevant infrastructure in place. If I ever get the time and interest to work on this again, I'll reach out to you.

Have you all considered adding scientific articles to your bibliographic database? Finding existing translations of scientific articles can be a real pain. I know because I spent a lot of time doing that during my PhD [1].

For a while I was collaborating with Victor Venema in the volunteer organization Translate Science [2] to try to create a bibliographic database of scientific translations, but unfortunately Victor died, and I became too busy to continue.

[1] https://academia.stackexchange.com/a/93209/31143

[2] https://translate-science.codeberg.page/

I think so too. If unclear, I don't use LLMs for coding at the moment and was just commenting on what I've seen from others who do in computational fluid dynamics.

Edit: Let me add that while I think it would be easy to instruct a LLM to do what I'd like, LLMs don't do these things by default despite them being recognized as best practices, and I'm not confident in LLMs getting the data or references right for validation tests. My own experience is that LLMs are pretty bad when it comes to reproducing citations, and they tend to miss a lot of the literature.

What I've observed in computational fluid dynamics is that LLMs seem to grab common validation cases used often in the literature, regardless of the relevance to the problem at hand. "Lid-driven cavity" cases were used by the two vibe coded simulators I commented on at r/cfd, for instance. I never liked the lid-driven cavity problem because it rarely ever resembles an actual use case. A way better validation case would be an experiment on the same type of problem the user intends to solve. I think the lid-driven cavity problem is often picked in the literature because the geometry is easy to set up, not because it's relevant or particularly challenging. I don't know if this problem is due to vibe coders not actually having a particular use case in mind or LLMs overemphasizing what's common.

LLMs seem to also avoid checking the math of the simulator. In CFD, this is called verification. The comparisons are almost exclusively against experiments (validation), but it's possible for a model to be implemented incorrectly and for calibration of the model to hide that fact. It's common to check the order-of-accuracy of the numerical scheme to test whether it was implemented correctly, but I haven't seen any vibe coders do that. (LLMs definitely know about that procedure as I've asked multiple LLMs about it before. It's not an obscure procedure.)