HN user

tobias2014

382 karma
Posts4
Comments127
View on HN

I agree, I see how Gemini itself with a usable harness can be excellent. But in agy with forced eager compaction (~125k with 3.1-pro, ~200k with flash) a kind of laziness and forgetting shows through that leads to an endless sequence of stopgap instructions, even with rigorous GEMINI.md and isolated task delegation and a good task tracking system. Agy is basically useless for more complex problems as far as I am concerned, at least when used somewhat autonomously as one could expect from claude. For strictly mechanical one-shot tasks it might be fine. I've spent way too much time working around these limitations instead of just continuing to use claude. Hoping that things would have improved with flash 3.6 I feel that it's actually worse in following instructions, and always acts even when just asked a question. If just agy offered a better experience and got rid of the terrible forced automatic eager compaction.

PS: That opus-4.6 via agy works so much better points in another direction though!

For any space-like event you can find reference frames where things happen in different order. For the time-like situation you described the order indeed exists within the cone, which is to say that causality exists.

GPT-5.2 7 months ago

I guess when you use it for generic "problem solving", brainstorming for solutions, this is great. That's what I use it for, and Gemini is my favorite model. I love when Gemini resists and suggests that I am wrong while explaining why. Either it's true, and I'm happy for that, or I can re-prompt based on the new information which doesn't allow for the mistake Gemini made.

On the other hand, I can also see why Claude is great for coding, for example. By default it is much more "structured". One can probably change these default personalities with some prompting, and many of the complaints found in this thread about either side are based on the assumption that you can use the same prompt for all models.

GPT-5.2 7 months ago

And who believes that the difference between 91.9% and 92.4% is significant in these benchmarks? Clearly these have margins of error that are swept under the rug.

How is "If" as a function even a drawback? It is largely seen as something desired, no? I would see that as a huge advantage, which allows for very powerful programming and meta-programming techniques.

This is not true. Mathematica has the concept of contexts. You can have each notebook have it's own unique context. Mathematica Packages create their own context too, we are not talking about module's here which are useful for local variable scoping. Packages and contexts lead to the isolation you are looking for. These are things that have been around since the initial Mathematica 1.0 in 1988 (!). https://reference.wolfram.com/language/ref/Context.html

Same about your criticism of error handling and control flow: https://reference.wolfram.com/language/guide/RobustnessAndEr...

I've been at a few universities and labs as a postdoc, and a Mathematica license always came either as part of the University or the department. It might not be relevant in some disciplines, but generally I assume it must be used a lot to warrant such broad licensing (it is a tool I use daily as a theoretical physicist).

Meanwhile companies exist that have built essentially layers in front of chatbots, masking or filtering sensitive data, then forwarding the masked query, then unmasking it when giving back to the user(e.g. https://www.liminal.ai/ ).

Ideally you shouldn't paste sensitive information into the chat in first place. But when such companies can guarantee certain compliance types, it might be better to offer this rather than letting people use chats uncontrolled in companies.

This might turn into a debate of defining "simplest", but I think the ensemble/statistical interpretation is really the most minimal in terms of fancy ideas or concepts like "wavefunction collapse" or "multiverses". It doesn't need a wavefunction collapse nor does it need multiverses.

M4 MacBook Pro 2 years ago

I hope you bring that up as an example in favor on open-source, as an example that open-source works. In a closed-source situation it would either not be detected or reach the light of day.

I was interested and looked at the actual patent: https://patents.justia.com/patent/4348422 (there seem to be multiple patent documents, but this one adds some explanation), and he writes "I have now surprisingly discovered".

https://tastydecafs.com/blogs/learn-about-decaf/co2-decaf further explains "The story of C02 decaffeination goes back to 1967. It was then when a chemist at Max Planck Institute named Kurt Zosel stumbled upon an interesting discovery. Zosel, like many other chemists, was using high-pressure C02 to remove individual substances from other mixtures."

It must have something to do with caffeine being an alkaloid, while coffee overall is acidic. So I suspect that this pressurized CO2 is able to dominantly remove such alkaloids... I leave the details to a chemist :)

This is why I think that modeling elementary physics is nothing else than fitting data. We might end up with something that we perceive as "simple", or not. But in any case all the fitting has been hidden in the process of ruling out models. It's just that a lot of the fitting process is (implicitly) being done by theorists; we come up with new models and that are then being falsified.

For example, how many parameters does the Standard Model have? It's not clear what you count as a parameter. Do you count the group structure, the other mathematical structure that has been "fitted" through decades of comparisons with experiments?

Recently talked to a DOE program manager. I confronted them with the question why we do things like machine learning or quantum computing in academia, when there's no hope in matching industry for those things. His answer was that from an official perspective this isn't the goal, but future workforce training for national interests and security. NSF might differ slightly, but I think this makes sense.

Is academia supposed to compete? I think the researchers wish for that, but that's not directly how government funding agencies see it. From a government's funding perspective the goal is to train tomorrow's workforce. People learn in academia and then transition and contribute outside. As far as agencies like DOE go, that is an explicit goal.

Is knowing the presence of something more valuable than knowing the precise absence of something? In my opinion both drive our knowledge forward and challenge our current understanding and development of models.

This is very interesting, because for me it's just the opposite. In particular the two column layout is just more readable and approachable for me. The PDF version also allows for a presentation just as the authors intended. I guess it's good that they offer both now.

"Tracking ratings variance by release year, we observe an increase in review variability starting in the late 1970s, followed by a precipitous decline in variance starting in the early 2000s."

I'm sure there is also some time dependent effect in how people rate movies now vs. in the past. For a snapshot in current time, one would need to be careful to not include ratings from too long ago as society and norms change.

Here are some good references on this with actual measurements, see e.g. table 4 on how many discharge cycles you can get out of a certain maximum voltage / charging level: https://batteryuniversity.com/article/bu-808-how-to-prolong-...

https://batteryuniversity.com/article/bu-204-how-do-lithium-... https://batteryuniversity.com/article/bu-501a-discharge-char...

On Thinkpads you can use tp_smapi to set charge start and stop thresholds https://wiki.archlinux.org/title/tp_smapi