HN user

amanda99

285 karma

Background in applied math, work in ML and software.

Posts2
Comments70
View on HN

I'm excited to see this. Have been using LiteLLM but it's honestly a huge mess once you peek under the hood, and it's being developed very iteratively and not very carefully. For example. for several months recently (haven't checked in ~a month though), their Ollama structured outputs were completely botched and just straight up broken. Docs are a hot mess, etc.

Human 1 year ago

I really liked the first chapter: the juxtaposition between ideas on how "AI"/algorithms/etc work and how humans work. I enjoyed how it was flipped around.

The second chapter was very disappointing and lost the intrigue I had built from the first.

Does this not require one to trust the hardware? I'm not an expert in hardware root of trust, etc, but if Intel (or whatever chip maker) decides to just sign code that doesn't do what they say it does (coerced or otherwise) or someone finds a vuln; would that not defeat the whole purpose?

I'm not entirely sure this is different than "security by contract", except the contracts get bigger and have more technology around them?

I think the OP's point here was that if it's a PR and it's ignored: you spent a bunch of time writing a PR (which may or may not have been valuable to you, e.g. if you maintain a fork now). On the other hand, if it was an esoteric contribution process, you spent a lot of time figuring out how to get the patch in there, but that obviously has 0 value outside contributing within that particular open source project.

I agree, and I also share your experience (guess I was a bit earlier with PHP).

I think what's left out though is that this is the experience of those who are really interested and for whom "it's not satisfying" to stay there.

As tech has turned into a money-maker, people aren't doing it for the satisfaction, they are doing it for the money. That appears to cause more corner cutting and less learning what's underneath instead of just doing the quickest fix that SO/LLM/whatever gives you.

Fun fact: in Gilbertese (the language of Kiribati), an "s" sound is written "ti". So Kiritimati is pronounced Kirismas (close to "Christmas").

On my journey there I met a fellow called "Simeon" and people would call him "Tim"...

A lot of folks are saying big ships are outdated. Keep in mind that war is 80% logistics: and if you are going to engage an enemy far away from your landmass (or project power for that matter), you need huge capacity to move materiel to said location, and that normally means ships.

An interesting thing to look at is something called the "Fully Burdened Cost of Fuel": for every gallon that the US delivers to a Forward Operating Base, they spend something like 6-20 gal getting it there.

The point is that you need to move insane amounts of stuff to fight a war effectively. The actual fighting is just the tip of an iceberg of logistics.

I'm sure the ad networks do a lot more than use high precision variables for soft-linking.

These are professional networks with a ton of capital thrown behind them. They have pretty decent algorithms, heuristics, etc; and you don't make money (compared to the other data correlation teams) if you do simple dumb stuff. I'm certain they take into account those trying to be privacy-conscious, if only to increase their match rates to be competitive.

They use TechSoup.org (whose site appears to be down rihgt now), but you can look up on that website information on who qualifies in Thailand. TS does the regional verification of what is a good enough not profit.

these numbers are just your perception.

Of course they are, I hoped it was clear I was just sharing my experience trying to use it for research!

I did in general word it as I would a question to a researcher, which includes an uncertainty in it being true. E.g. this is from a recent prompt: "is this true in general, if not, what are the conditions for this to be true?"

I think you are imagining a different class of "questions".

To clarify, I was doing research on applied math. My field is not analysis, but I needed to prove some bounds on certain messed up expressions (involving special functions, etc), and analyze an ODE that's not analytically solvable. I used the COT model a fair bit.

I would ask ChatGPT for hints/ideas/direction in proving various bounds, asking it for theorems or similar results in literature. This is exactly the kind of thing where a researcher would go "yeah this looks like X" or "I think I saw something like this in (book/article name)", or just know a method; or alternatively say they have no clue. ChatGPT most often will confidently give me a "solution", being right 10% of the time (when there's a pretty standard way to do it that I didn't see/know).

On the whole it was quite useful.

AI is often wrong, never knows when it's wrong, but people are like this too.

When talking with various models of ChatGPT about research math, my biggest gripe is that it's either confidently right (10% of my work) or confidently wrong (90%). A human researcher would be right 15% of the time, unsure 50% of the time, and give helpful ideas that are right/helpful (25%) or wrong/a red herring (10%). And only 5% of the time would a good researcher be confidently wrong in a way that ChatGPT is often.

In other words, ChatGPT completely lacks the meta-layer of "having a feeling/knowing how confident it is", which is so useful in research.

Is this not kind of trivial, at least the way they've defined it? It's kind of very obviously NP complete to me. Any of these text problems where you're tying to optimize over a whole corpus is kind of not hard to see to be NP-complete, similar to longest common subsequence.

I thought this was going to be about clean cooking fuels. One of the significant projects of the WHO is transitioning the world towards cooking with clean fuels that reduce indoor air pollution. In the worst case certain populations are burning plastics to heat their water, food, and homes, and as you can imagine this is incredibly destructive to health.