HN user

sireat

2,852 karma

This handle is unique to HN. I do not use it anywhere else. You could probably find me by finding similar posts by someone else on Reddit and maybe just maybe on Usenet.

Posts2
Comments1,605
View on HN

This is rather like SAT from 35 years ago.

Same strategies apply for guessing the unknown especially with a modicum(it was on the test!) of Latin knowledge..

Strange that pretty every one here is getting 70k estimates (93/100 for me).

Feels a bit high at least for me as a non-native speaker.

I got 2 words I knew wrong, and guessed about 5 unknown words correctly. Those were bizarre repetitive words I've never seen before.

I remember doing a similar test from a reputable university about 10-15 years ago also in an app format and only got about 30k estimate.

As long as Codex remains so affordable and useful they do not have to slash prices, just keep Codex usable.

I keep meaning to try Claude Code, but I can't seem to run out of limits on Codex on regular pro plan.

Meanwhile all my friends on Claude Code are fighting the token limits every few hours.

I even switched to using extra high for easy medium level script tasks as a test and besides taking longer there was not much reduction in the token allowance.

I generally write a detailed spec before plan then possibly iterate a bit before implementation. Not sure what I am doing "wrong".

Fantastic practical achievement!

I wonder if I could get similar or even better performance from similar Dell T7610 workstation with dual Xeons and also 128GB DDR3?

The CPUs are better core wise, but that probably does not make much difference?

It has CPUs 2 × Xeon E5-2697 v2

Cores / threads 24 cores / 48 threads total

Per-CPU cores 12 cores / 24 threads

Base clock 2.70 GHz

Max turbo 3.50 GHz

It is sitting gather dust but reading spead Gemma sounds promising.

I am paraphrasing but I think it was W. Buffett who said:

"Work at the job that you do not hate"

In other words, not all vocations that you are great at and talented and want to pursue are valued by current world.

I love playing chess way more and actually am reasonably good at it, but programming and teaching are valued more and I like those too.

As Jimmy O. Yang's father reportedly said: "Pursuing your dreams is how you become homeless"

https://www.youtube.com/watch?v=GO6ntvIwT2k&t=22s

At the same time you have to be out there in the world, increase your luck surface - if you sit in your cubicle/room/private chatroom all day you are less likely to make a mark on the world despite your brilliance.

Again I forgot which artist said it but that in New York art scene the most successful artists spent most of their working days socializing not painting/sculpting etc.

Along with already mentioned "Buy, Borrow, Die" strategy is the more widely practiced "Expense everything" strategy which often ends in tax disaster for the practitioner.

One of the early adopters was https://en.wikipedia.org/wiki/B._C._Forbes the founder of Forbes.

He expensed lavish Gatsby style parties and everything.

I remember reading a biography of his that one way in 1920s he accomplished was by having bought some big mostly useless plot of land and technically his lavish parties were sales presentations to sell this land. Occasionally some of his acquintances would actually buy a parcel of mostly useless land in middle of nowhere thus the business use was actually maintained. Again, highly unlikely to fly today with IRS and even then there were tax lawsuits.

The issue is that it is impossibly hard to pull off without going into tax fraud territory.

Another interesting case of "Expense everything" were ABBAs stage dresses and suits. They were purposely flashily impractical to avoid falling afoul of Swedish tax laws.

That said tax authorities in most countries do allow some leeway for the small fish. Basically pragmatic tax authorities give you certain limits for certain expenses that you can expense.

So in my European country you can expense a certain amount of gas, travel, clothing, eating out, etc as a self-employed. Yes you should have receipts, but if you stay within limits, it is up to you how honest you want to be about that "business" lunch.

I remember it being it common in US too, someone takes you to lunch and you are supposed to mention their business and talk a few minutes about their business, then in their eyes it was a business expense.

However, the moment you start going over these limits you will face increased scrutiny and you are in for a bad time for claiming as business expense lunch with your friends at Dorsia.

The latest clickbait style can be mitigated by custom instructions. I use: "Tell it like it is; don't sugar-coat responses. Use academic university level explanations unless instructed otherwise. Do not end with teaser offers or curiosity hooks. Give the full answer immediately. If related topics exist, show them as a brief bullet list. Use professional language and style."

Now I actually often like the related topics hooks, just not the clickbaity version from last few weeks.

If not for Codex performing so well for me from VS Code I'd happily migrate to Claude or Gemini.

Basically you have Cremant type sparking wines which are produced from other regions of France besides Champagne. It is just like Champagne just that other French regions like Loire, Alsace, Bordeux etc are not allowed to call it Champagne.

So just like Armanac's are like Cognac's for lower price, good Cremant will be cheaper and more enjoyable that cheaper Champagne (I've not had any really expensive Champagne).

Then you have Cava from Spain which is similar process to Cremants and Champagne. The difference would be in type of grapes used. A friend of mine swears by Cavas just like I swear by Cremants from Loire region. However my wife hates Cava.

Then Proseccos from Italy again are similar, but quality varies more.

After that we get into more questionable cheaper sparkling wines which usually means some sort of out of bottle insertion of CO2 and even worse version include some other modifications such as sugar.

In general to avoid literal headaches you want BRUTs. Anything semi-sweet or sweet is suspicous.

Again I am not a full wine expert but this is mostly years of ahem experience.

After, RAM, SSD, GPUs, now HDDs what else is there left to sell out? Power supplies, fans?

In a way this feels a bit absurd for these AI centers to hog HDDs.

As pointed by others neither training nor inference require HDDs and storing raw data should not require that much.

So my hypothesis is that it is a double whammy of overall declining consumer sided HDD demand, leaving data centers as main source of demand and additional demand from the new AI centers.

I feel like the AI centers are just buying HDDs because why not throw a HDD in each server blade even if there is no need? The money is there to be spent and it must be spent.

As someone who has been building computers since 1989 it feels like end of personal hobby casual building.

I will end with an imperfect analogy with multiplayer gaming. It is quite common in multiplayer games for higher level players to wish to acquire some tradeskill they neglected to acquire earlier. maybe a new quest appears, or new "must have" item that requires such skill.

They (past me included) have too much game money and no wish to acquire tradeskill items slowly. So the "rich" will overpay by 2x or 10x or even 100x the usual price.

That is free market at work right?

In the process whole low level economy is destroyed due to 2nd order effects. Meaning a new player starting out can only be a farmer.

So if a student comes to me wishing to start building computers what advice do I give them? Farm something?

This is very cool and having stalemate is nice, however how much space would it take to implement the full ruleset?

As you write: not implemented: castling, en passant, promotion, repetition, 50-move rule - those are all required to call the game being played modern chess.

I could see an argument for skipping repetition and 50-move rule for tiny engines, but you do need castling, en pessant and promotion for pretty much any serious play.

https://en.wikipedia.org/wiki/Video_Chess fit in 4k and supported fuller ruleset in 1980 did it not?

So I would ask what is the smallest fully UCI (https://www.chessprogramming.org/UCI) compliant engine available currently?

This would be a fun goal to beat - make something tiny that supports full ruleset.

PS my first chess computer in early 1980s was this: https://www.ismenio.com/chess_fidelity_cc3.html - it also supported castling, en pessant, not sure about 50 move rule.

Interesting information but these are not hard numbers.

Surely the 100-char string information of 141 bytes is not correct as it would only apply to ASCII 100-char strings.

It would be more useful to know the overhead for unicode strings presumably utf-8 encoded. And again I would presume 100-Emoji string would take 441 bytes (just a hypothesis) and 100-umlaut chars string would take 241bytes.

Very simple - look for who has a stake in Groq currently:

https://www.cnbc.com/2025/12/24/nvidia-buying-ai-chip-startu...

"Davis, whose firm has invested more than half a billion dollars in Groq since the company was founded in 2016, said the deal came together quickly. Groq raised $750 million at a valuation of about $6.9 billion three months ago. Investors in the round included Blackrock and Neuberger Berman, as well as Samsung, Cisco , Altimeter and 1789 Capital, where Donald Trump Jr. is a partner."

POP QUIZ - Which minority partner is the key here?

LLM Year in Review 7 months ago

What is current state of the art workflow when working with legacy code across multiple languages?

This would be a 100 kLOC legacy project written in C++, Python, and jQuery era Javascript circa 2010. Original devs have long left. I would rather avoid C++ as much as possible.

I've been Github Copilot (in VS Code) user since June of 2021 and still use it heavily, but the "more powerful intellisence" approach is limiting me on legacy projects.

Presumably I need to provide more context on larger projects.

I can get pretty far with just ChatGPT plus and feeding bits and pieces of project. However that seems like using the wrong tool.

Codex seems better for building things but not sure about grokking existing things.

Would Cursor be more suitable for just dumping the whole project (all languages) basically 4 different sub projects and then selectively activating what to include in queries?

How about picture gen though on Dall-E?

It is so infuriating to get content block on ChatGPT for pretty much any fairy tale that has had a Disney related adaptation.

Try getting a Grimm's 19th century Snow White illustrations. You can not because the Disney crap supersedes it.

In fact you can not get a Snow White illustration of any kind on ChatGPT.

I can not figure out any prompts that would draw using public domain knowledge.

Same goes for a pirate fighting a flying boy - no good.

New one this week was when I tried to draw a border around my daughter's picture of a Poppy from Trolls(That's Dreamworks but same problem).

The actual copyrighted Poppy appeared in the border half way down the generation and then of course content block appeared.

What is hilarious though that ChatGPT will profusely apologize and provide extremely detailed instructions in setting up local Stable Diffusion as an alternative...

As I recall Holmes did in fact do a lot of walking. He vacillated between periods of inactivity(cocaine, violin, shooting V in wall with a revolver) and intense activity (taking up disguises and doing various physical activities including walking all across London and elsewhere.

Just because your logical mind says one thing is good to do and you know you should do it you are not going to always obey your rider, the inertia of the elephant takes over.

So you need a trigger to snap out of it, for Holmes it was a new case.

Indeed regular Jupyter works so well on VS Code for solo work these days that there is no real need for a new entrant.

So what pain point are these new entrants trying to solve?

Sure there is an issue of .ipynb basically being a gnarly json ill suited for git but it is rare that I need to track down a particular git commit. Even then that json is not that hard to read.

Also I'd like an easier way to copy cells across different Jupyter notebooks, but at the end of day it is just Python and markdown not very hard to grok.

OpenAI has ridiculous guardrails for illustrations covering any public domain subject that has been covered by Disney or any other major public corporation.

So by that benchmark Japanese companies have a case.

Try generating a 19th century style illustration of Snow White. You can't at least not on OpenAI platform.

Try generating a picture "of flying boy fighting a pirate on a ship".

I have a story from the mechanical side.

I spent a month in 2012 roughly 4 hours a day doing various tasks.

It was horrible, even if I followed all the "best practices" of Turkers it was not a way to make a living.

By end of the month, I had become so jaded to all the "priming" experiments by graduate and undergraduate psychology students. Those usually paid at least something 3-4 USD an hour.

Did some porn labeling tasks, those were horrible after the novelty wore off.

Did very few other labeling tasks because they paid next to nothing.

To have someone actually depend on living for these seemed like a torture.

This 30 Euro jump in Europe was a kick in the pants for me.

Even though it is still a relatively good deal for a Family Plan (compared to say Google Drive or Dropbox) for OneDrive, I finally dropped my Microsoft 365 Family plan.

The final straw was that the Copilot was completely unhelpful and hallucinated features Office portal does not have.

Not OP, but CLIP from OpenAi (2021) seems pretty standard and gives great results at least in English (not so good in rarer languages).

https://opencv.org/blog/clip/

Essentially CLIP lets to encode both text and images in same vector space.

It is really easy and pretty fast too generate embeddings. Took less than hour on Google Colab.

I made a quick and dirty Flask app that lets me query my own collection of pictures and provide most relevant ones via cosine similarity.

You can query pretty much anything on CLIP (metaphors, lightning, object, time, location etc).

From what I understand many photo apps offer CLIP embedding search these days including Immich - https://meichthys.github.io/foss_photo_libraries/

Alternatives could be something like BLIP.

Like Simon I've started to use camera for random ChatGPT research. For one ChatGPT works fantastically at random bird identification (along with pretty much all other features and likely location) - https://xkcd.com/1425/

There is one big failure mode though - ChatGPT hallucinates middle of simple textual OCR tasks!

I will feed ChatGPT a simple computer hardware invoice with 10 items - out comes perfect first few items, then likely but fake middle items (like MSI 4060 16GB instead of Asus 5060 Ti 16GB) and last few items are again correct.

If you start prompting with hints, the model will keep making up other models and manufacturers, it will apologize and come up with incorrect Gigabyte 5070.

I can forgive mistaking 5060 for 5080 - see https://www.theguardian.com/books/booksblog/2014/may/01/scan... . However how can the model completely misread the manufacturers??

This would be trivially fixed by reverting to Tesseract based models like ChatGPT used to do.

PS Just tried it again and 3rd item instead of correct GSKILL it gave Kingston as manufacturer for RAM.

Basically ChatGPT sort of OCRs like a human would, by scanning first then sort of confabulating middle and then getting the footer correct.

To be a bit flippant, you can absolutely destroy energy by creating some mass..

Then again most of us do not have particle accelerator nearby looking for Higgs boson.

While I do have multiple OpenRouter accounts(personal and organizational) I did not even look into concurrent calls - it was sequential.

The job was set on Friday and ready on Monday. On average it was about 5k tokens (documents ranging from 1k to 200k in size) and only about 10 tokens out.

Average response was about 1.5 seconds ~ 40 hours for full set.

I really did some heavy prompt testing to limit output.

Even then every few thousand queries you'd get some double token responses. That is Gemini would respond in duplicate - ie Daisy Daisy.

Basically it boils down that for most queries google/gemini-2.5-flash is the workhorse fast/cheap/good enough.

Add in multimodality, 1M context and it is such a Swiss army knife.

It is cheap and performant enough to run 100k queries. (Took a bit over a day and cost around 30 Euros for a major document classification task). Yes in theory this could have been done with fine-tuned BERT or maybe even with some older methods but it saved way too much time.

There is another factor that may explain why Flash is #1 in most categories on OpenRouter - Flash has gotten reasonably decent at less common human languages.

Most cheap (including Flash Lite) and local models mostly have English focused training.

Curious, what kind of prompt gives you the same text adventure game?

Surely it is a question of prompting some context(in UI mode) or with additional kicker of temperature (if using API)?

At the very least some set up prompt such as "Give me 5 scenarios for text adventure game" would break the sameness?

There have always been theories that OpenAI and other LLM providers cache some responses - this could be one hypothesis.

Fun but its Python REPL broke the immersion for me.. Python 3.13.2 (main, Aug 4 2025 20:25:58)

I was expecting Python 2.2 or 2.3 ... not sure what was the earliest version of Python on Pyodide