HN user

bugglebeetle

1,518 karma
Posts0
Comments1,150
View on HN
No posts found.

It’s not “deeply weird” but was intentionally legislated to prevent labor from obtaining similar sectorial bargaining power and political influence as they had in Europe. Some entertainment industry unions had already achieved similar stature pre-WWII and were grandfathered in, but the US learned their lesson here and explicitly sought to atomize at the workplace level.

Unfortunately, this looks to only cover the larger MoE models. I imagine the smaller models are what most people would target. 9B just dropped two days ago, so not surprised it’s not explicitly documented, but does use a hybrid mamba architecture that I expect needs some special consideration.

Ireland’s affordability problems are almost exclusively centered around its housing crisis and they need to just commit themselves to over-supply induced wealth destruction for the landlord class and older generations. Thankfully, there demographics also support such a move.

I tested this pretty extensively and it has a common failure mode that prevents me from using: extracting footnotes and similar from the full text of academic works. For some reason, many of these models are trained in a way that results in these being excluded, despite these document sections often containing import details and context. Both versions of DeepseekOCR have the same problem. Of the others I’ve tested, dot-ocr in layout mode works best (but is slow) and then datalab’s chandra model (which is larger and has bad license constraints).

Everyone who tells the story of the reformation leaves out that Martin Luther also used this new technology to widely disseminate his deranged anti-Semitic lies and conspiracies, leading to pogroms against Jews, a hundred years of war across Europe, and providing the ideological basis for the rise of Nazism.

Nah, I don’t miss at all typing all the tests, CLIs, and APIs I’ve created hundreds of times before. I dunno if I it’s because I do ML stuff, but it’s almost all “think a lot about something, do some math, and and then type thousands of lines of the same stuff around the interesting work.”

LLMs Are Not Fun 7 months ago

I just had Claude Code finetune a reranker model to improve it significantly across a large set of evals. I chose the model to fine tune, the loss function, created the underlying training dataset for the re-ranking task, and designed the evals. What thinking did I outsource exactly?

I guess did not waste time learning the failure-prone arcana of how to schedule training jobs on HuggingFace, but that also seems to me like a net benefit.

LLMs Are Not Fun 7 months ago

As a former artist, I can tell you that you will never have good or sufficient ideas for your art or writing if you don’t do your laundry and dishes.

A good proxy for understanding this reality is that wealthy people who pay people to do all of these things for them have almost uniformly terrible ideas. This is even true for artists themselves. Have you ever noticed how that the albums all tend to get worse the more successful the musicians become?

It’s mundanity and tedium that forces your mind to reach out for more creative things and when you subtract that completely from your life, you’re generally left with self-indulgence instead of hunger.

LLMs Are Not Fun 7 months ago

For me, the joy of programming is understanding a problem in full depth, so that when considering a change, I can follow the ripples through the connected components of the system.

The joy of management is seeing my colleagues learn and excel, carving their own paths as they grow. Watching them rise to new challenges. As they grow, I learn from their growth; mentoring benefits the mentor alongside the mentee.

I fail to grasp how using LLMs precludes either of these things. If anything, doing so allows me to more quickly navigate and understand codebases. I can immediately ask questions or check my assumptions against anything I encounter.

Likewise, I don’t find myself doing less mentorship, but focusing that on higher-level guidance. It’s great that, for example, I can tell a junior to use Claude to explore X,Y, or Z design pattern and they can get their own questions answered beyond the limited scope of my time. I remember seniors being dicks to me in my early career because they were overworked or thought my questions were beneath them. Now, no one really has to encounter stuff like that if they don’t want to.

I’m not even the most AI-pilled person I know or on my team, but it just seems so staggeringly obvious how much of a force multiplier this stuff has become over the last 3-6 months.

Almost nothing aside from children’s books is written exclusively in hiragana or katakana. You have to also memorize the variable readings of about 2000 kanji and many texts are nearly unintelligible without them. Pretty much everyone can memorize the former, but must struggle with the latter.

Both Korean and Mandarin are simpler in this regard (and the latter follows the same grammatical order as English).