HN user

berkes

6,038 karma

Ruby and Rust Developer.

[ my public key: https://keybase.io/berkes; my proof: https://keybase.io/berkes/sigs/uXC3DKS9jat3AILQSm_fIb03uxFV1JSKz8MqCKwzD90 ]

Posts26
Comments2,437
View on HN
www.youtube.com 8mo ago

Europe's Plan to Unleash Its Startups| 28th Regime

berkes
1pts0
love.berk.es 1y ago

Show HN: A unique generated maze to share with your Valentine

berkes
100pts20
mastodon.social 2y ago

Zuck threads.net: First post in the Fediverse

berkes
2pts0
www.forbes.com 3y ago

Batteries and Renewables Are Saving Texas During the Heat Wave

berkes
78pts98
www.youtube.com 3y ago

The Dumbest Excuse for Bad Cities

berkes
4pts0
search.flockingbird.social 3y ago

Show HN: Search Job Openings on the Fediverse

berkes
3pts1
brainbaking.com 3y ago

My Desktop Is Dull Thanks to macOS

berkes
2pts0
berk.es 3y ago

“How do I test X” is almost always answered with “by controlling X

berkes
1pts0
berk.es 3y ago

Exponential compound interest on Technical Debt. And how I avoided it

berkes
1pts0
berk.es 3y ago

Event Source Your Spreadsheets for Flexibility and Maintainability

berkes
2pts1
strapi.io 4y ago

Strapi closes $31M Series B led by CRV

berkes
1pts0
colinmorris.github.io 4y ago

In search of the least viewed article on Wikipedia

berkes
1pts1
berk.es 4y ago

The Waning of Ruby and Rails

berkes
9pts5
news.ycombinator.com 4y ago

Ask HN: Was Outline.com Taken Offline?

berkes
5pts4
www.reddit.com 5y ago

I've tested Rust HTTP clients again, found over 50 bugs

berkes
2pts2
investors.etsy.com 5y ago

Etsy to acquire global fashion resale marketplace Depop

berkes
2pts0
news.ycombinator.com 6y ago

Ask HN: What podcasts do you recommend today?

berkes
2pts2
blog.joinmastodon.org 6y ago

Why ActivityPub Is the Future

berkes
7pts8
lwn.net 6y ago

An uproar over the Fedora Git forge decision

berkes
4pts2
ipfs.io 7y ago

Mirror of OpenStreetMap databasefiles on IPFS

berkes
3pts4
berk.es 7y ago

The Private Blockchain Fallacy

berkes
110pts107
github.com 8y ago

Microsoft: Drop ICE

berkes
5pts0
berk.es 8y ago

Making a case for JavaScript, in-browser Mining

berkes
45pts74
news.ycombinator.com 8y ago

Ask HN: Has anyone built an API or Service that handles invitation workflows

berkes
1pts0
planetruby.herokuapp.com 11y ago

Planet Ruby

berkes
51pts8
acquia.com 15y ago

Twitter Selects Drupal to Power Developer Community Website

berkes
1pts0

And also, there's certainly places and situations where you do want to spend tens of thousands on PLCs, buttons and wiring.

But not all of it needs this level, quality or insurance. A factory, farm, supermarket, stage, etc, is filled with tech. Many of it critical for operation, but just as many is non-critical. And either this "non critical" must then be acquired with the same level of quality, insurance and costs as the critical. Or it's simply not available - likely because there's no demand for it at this price point.

be able to survive field conditions

you are right. But my point was that not all of needs to survive the same conditions.

E.g. in a theater, you want the high class industry grade rigs for the main lightning when you are running the shows. But during rehearsals, or during the development of the stage (which is somewhat iterative by nature), custom made stuff suffices.

And this is where currently, custom made stuff is deployed. Stage technicians use a remarkable lot of intermediate, custom made hard and software. Some 3d printed valve to find out if the idea of the valve works at all, before ordering a CNC aluminium version. Some arduino before ordering that €2k/month controller to find out if the separate controller is needed at all.

This allows a show to prototype and experiment before they order the €20k water-spraying system. Or allows them to rehearse for months without paying €5k a month for the high class lightning rigs and only rent those for the actual setup.

Its the same with my other examples. They fill a niche. They don't replace the entire existing market.

the average farmer or show operator doesn't have the skill

Exactly! That's why there's an opportunity here. "We" software engineers tend to build a rediculous amount of SAAS that are aimed at other software-engineers, because that's were we notice the problems and opportunities. My advice always, is to look beyond your bubble. The examples above are real-life examples where currently "farmers", "show operators" and "fitter" already are cludging together solutions to solve their problems. IMO there's no better proof that this is a problem to research as potential tech startup, than non-techies cludging together tech to solve their problems.

There are many opportunities to retrofit old systems of all types with modern low-cost embedded technologies.

I'd leave out the "old" there.

An example: Friends work(ed) as engineers in theaters and music-venues. The costs for (DMX) equipment is insane. Some lamps, rigs or controllers are so expensive that shows only rent them. While there's lots of cases where you don't want to cut costs¹, quite often a cheap variant cuts it just fine. So apparently there's a lot of Pie's, Arduinos, ESP32s and old androids used to handle stuff. Custom, hacked, duckt-taped solutions. All of them are certain a startup offering generic but cheap modules that engineers can customize, would be an instant hit. Especially if "service" can be bought with it.

Or a friend who installs (climate) control systems in (food)factories, who cludges together arduino's and some hacky web-interfaces, who is certain that the industry would love some "off the shelve" solutions that installers can build their own hardware- and user-interfaces on is a very much wanted product. The industry (in EU) is according to him, one built on "lock ins", where a factory starts off with AcmeInc controllers and then slowly Acme will start charging almost "extortion-fee" alike figures just to fix stuff on a sunday night.

Or, a friend who has a fruit farm with lot's of freezer-storage (running hundreds of KWs), whith whom I cludged together some Arduino's and home-automation to control the freezers to run when electricity is cheap or free. Stuff like this is available for consumers at reasonable prices, but the moment the Amperages go to industry-level, the costs explode. Where a simple remote controllable power-switch can cost €1000s. I'm certain that there's hundreds of thousands of small warehouses, supermarkets, traders, logistics that could save a few hundreds to thousands of €'s a year, by controlling their freezers, heaters, pumps, pressurizers, dryers and whatnots smarter with "turn key" modules built from generic low cost hardware.

--

¹ relevant: There's a lot of "Nobody Ever Got Fired for Buying IBM" behind it: if you go for the cheap €800 chinese knock-off and mid-show something starts failing, its your fault. But if that €10.000 controller starts failing, it's no-ones fault.

I've worked on a complex bookkeeping SAAS. Our "CVS" im- and export was the most used feature by far.

Yet it was implemented on a thursday afternoon by a junior, using the first "library" that popped up in a google search. And over the years, leaked into every corner of the application. Unaffordable technical debt.

I cemented this experience into memory and now, in every new gig where I have to do "CSV", I isolate it, abstract it, overengineer it. Hell, last year I even built and launched a dedicated e2e tested csv-export service instead of just `from csv import writer` and call it a day.

Because CSV is the "interface" that will bring your system down if not properly designed.

Give that engineer a raise and an official title of "chaos monkey".

On the scale of AWS, you want your production system to be poked at and pushed against. Obviously not something to implement out of the blue in your production if you've never had it. But certainly worth considering.

Even on a smaller scale, it makes a lot of sense to have some deamon randomly bringing nodes down, shut off a database, fill a disk up, etc etc. In production.

Because that brings awareness to all engineers that the stuff that "can happen", will happen. It makes for a culture of defensive, considerate engineering. A culture where someone who finds bugs or holes in production is rewarded, not punished.

A culture where having some automated management for nodes is not a ticket "at the bottom the backlog", but a crucial attribute when picking or engineering the hosting system. That having redundancy in critical services like a database, is not an expense, but actually a cost saver, because chance of critical services failing is 100%, rather then some made up gut-feeling-chance. That the error handling and tests for when a module cannot write a file to disk isn't something that "this senior who left 8 months ago overengineered", but a pattern to be consistently implemented¹. etc.

---

¹ It is one the things I love about rust, that it has built in enforcement for you to deal with all the esoteric errors (e.g. https://doc.rust-lang.org/std/io/enum.ErrorKind.html#variant...) that can occur. Sure, you can write `File::create("foo.txt").unwrap()` but even then, you have made a conscious decision, clearly "documented" that you will crash the program - i.e. postponed the decisions on what (groups of) errors must be handled in what way, to a time you have the information to make that decision. Same design decision with go, where errors are explicitly to be handled by the programmer.

The service needs to bill "credits"

When talking about e2e tests, it matters where the e's are. If they are at the public interface of "the service", then, indeed, you'd test against "credits".

But ideally, there'd be e2e tests that test the consumers' experience, where the 'e's are the web interface, emails, CLI and other things consumers click on, read, enter etc. This is where a "test currency" makes sense.

Almost all of them have one number that appears four times, and one or two that appear three times

To me that looked suspiciously like string-handling in a weakly typed language.

Like when you do `"100" + 1` in JavaScript, or `int("100" * 2)` in Python.

I've seen my share of such bugs in PHP, Python, Ruby, JavaScript. In production. Obviously not as simple as the examples, but subtle, like when a library update changed `someFancyLocalStorage.getOrDefault("lastOrder", 100)` by always casting the value to the type of the default (released as patch release). Or where typedEnvGet() should typecast "numbers", but keeps it a string when theres whitespace `AMOUNT_PER_CALL=100\n`. Or where a number passes through a deep stack of middleware and 99.9% of the times remains an int but in rare race conditions becomes a string. etc.

No evidence that's the case here. But from my experience, the repeating and strange formats of numbers hint strongly in that direction.

I was merely re-iterating the marketing from Apple on why they charge the 30%.

I personally think it's a lie at best, or a scam (on multinational level) at worst. I've always been convinced that regulation should force companies that make such statements to be held responsible. Treat them like an insurance company or such.

I pay a few dollars every month to be protected against scams and malware. So when I do get infected, Apple failed to do what I paid it for and Apple should be forced to pay me for damages.

Down to the manufacturer of the whole product you're buying.

In case of an app, what is the "product" you are buying? Because according to Apple, they add a lot of "value" by ensuring the software is safe, performant, etc etc. Am I not buying "a safe, checked app"? Or am I buying an app and then separately pay Apple for an added service of "checking the app for safety" etc etc. I'd very much presume the first.

But if its separate, "an app" can be rather ambigous. For a one-time-purchase game, its clear. But many apps are really a service or even more that happen to have "an app" as one of the ways to interact with the service: Netflix, Uber, protonmail, Vinted (or ebay), etc etc: the app isn't the thing I buy. It's a wrapper around a service. Or even just one of the portals through which I can buy stuff. Point being: It's not simple, so your answers don't fit the analogy of "wallmart".

I find this hard to believe, since Google Play store requires much less info, and doesn't disclose it to consumers (AFAIK).

Did the EU specifically demand this from Apple? Did they specifically require that consumers must be able to contact developers?

Or is this another "spin" by Apple to make the EU look bad when it imposes consumer protection that is bad for Apples revenue? Like they did with "chargers", "cables" and like the ad- and surveillance-industry has done quite successfully with their "spin" on the GDPR (making it seem like the EU or GDPR requires cookie banners - which it doesnt)

due to your political bias.

That's a wild accusation. Unless you happen to know me, but I don't think you do.

"Lobbying to reduce taxes" was a part of an example. The big problem is their "lobbying for greater freedom".

Yes, governments should be held in check. Systems strengthened to keep governments afraid of the people (not companies) they govern. So that they can keep those that are stronger than me and use that strength to harm me, amongst wich are companies, in check. We have most of this in place. It can (and should be) improved. But we have very little in place to keep companies in check. Antitrust cannot (or is not) used against monopolies, (so the common free-market mechanisms of peoples free movement between suppliers aren't working anymore). Laws that protect humans against companies abolished or not modernized. Information (on which agents in the free-market must operate) is less and less trustworthy, mostly a result of companies spreading that (mis)information.

And TBC, companies have harmed people, are doing so as we speak, whole cities, whole countries. They have actively destabilized governments. Not just "somewhere in Africa", but in the US, in Europe. That's an undeniable fact. Many companies will then spend a lot of effort (and money and influence) to be not held accountable (that is harder to prove, obviously). Many companies will wield influence to be allowed to wield firepower, in all it's forms and shapes.

I am from a country with one of the worst examples of such a company in human history, the V.O.C. Who had such power that it had private armies, was accountable to no-one, grew so large that no "official government" could force it to do anything, for over a century. It literally enslaved people, killed inefficient workers, over-exploit and deplete entire ecosystems, etc etc. We should actively and strongly work against such companies from ever coming back (if they ever really left)

Read more history.

That government has the monopoly of violence against the company too. They can easily strongarm the company into harming me.

With private security companies, private jails, private utilities, the harm both can and will do, is practically the same. And with the boundaries of "private" and "government" becoming more opaque (like private tech companies deciding who is to be detained and who not; who is allowed to access their money; or do their jobs, etc) the difference is even more theoretic.

A government that has absolute power over a company/sector, can (and will, in my belief) use that power to make the company/sector do it's work. Making this company/sector just as dangerous as the company.

If I get mistreated at an airport (of a democratic country), and the agents doing the mistreating are civil-servants, they (or their chain of command) are democratically held accountable. But if they are employees of Acme-Security Inc, granted the right of violence by the democratic government, there's nothing I can do, no-one to appeal to. Both cases are something I, an individual, cannot do anything against. I am mistreated anyway.

If I need to pay someone 300k to make the model and infrastructure

I was arguing for the existing AI-companies that already make and offer niche models. Like Mistral. But AFAICS, all AI companies have and offer such models.

So all you need to do, is use the existing models. And, yes, select it.

Which, ironically, I would highly value as a niche model myself. I spend way too much time following the breakneck race of the various companies just to pick the right models for my tasks at hand. "Just pick the latest" often yields worse results, or is magnitudes more expensive, or significantly slower, or all of it. "Just pick the most popular" can prove expensive, inefficient for some task etc. This investment, also ironically, has proven something of a "moat" for me. I know very well what Mistral and Anthropic offer. So I won't even bother with OpenAI, Google, X, Tencent etc etc etc models. I just don't have the time to keep researching the latest offers for their pros and cons.

A model that acts as "decision maker" and as proxy, as a conductor, that directs and transforms my questions and sends them off to the right model, right tools, right MCP etc, would be very welcome for me. So that I can just pick that one, and have the highly dynamic world of LLMs and other models shift like undercurrent beneath the surface of this One Model To Rule Them all.

The General Models' business-model is also looking more weak every iteration.

Costs of simple tasks grow extensively: OCR with "Mistral OCR" at $4 per 1000 pages vs OCR with Opus 4.8 at sometimes¹ $1 per "page".

Or just the immense costs when burning tokens in an unoptimized agentic coding environment costing tens of dollars for a few simple classes or functions versus a highly optimized "autocomplete" model costing under $10 for thousands of such classes and functions.

Or the, over ten dollars worth of tokens when some "agent" using a general model, tries to perform the task I gave it to "read the event on example.com/event/1337 and put it in my calendar", include commute time as well"

The "general models" currently only become smarter by growing bigger and having larger context windows - by becoming exponentially more expensive to train and to run and to interact with. Whereas "Niche" models can do the things that "normal code" cannot do, and improve by tuning and tweaking only that. Their goal is then to fill in gaps that traditionally are hard or impossible with normal software. Wheras the goal of a general model (with agentic reasoning)is to replace that entire "normal software".

One example: I am not interested in "chatting with my calendar". I'm interested in a calendar because it is a well known view (UI) of my planning and tasks, but I see a lot of opportunities where AI can improve my working with this calendar. I may be interested in a smarter screen when I hit "+ Add event"; one that has knowledge of my previous events and patterns (some RAG vector db maybe). One that maybe has access to content I just copied, or read (though: privacy?) or can open my camera to let me shoot a pic of something that has the event info on it. In such a set-up, Niche LLMs perform dedicated tasks: determine patterns (he always books a Yoga class on wednesday or thursday, two days in advance, so lets suggest a yoga class), determine existing content (event is planned 100Km from his home, so lets suggest the commute based on previous commutes like this). Or an OCR model. Or an autocomplete model. Relatively simple, niche models, called from within software to aid me when "calendaring". Not replace the entire calendar with some chat.

¹Edit: This was a rather unscientific research of mine, where I compared some models to read from photographs, compared purely on costs and timing. "Opus" or other generic LLMS with image input capabilities commonly did better on "performance" esp with difficult input such as a picture of a poster of some rock event.

Companies fall under the government. So what a company harvests (to sell more toilet bowl cleaner), is accessible to the government it falls under.

By that logic, you should fear companies at least as much as their governments when handing them your data.

But companies have additional goals: to increase profit. Which can be achieved by selling more toilet bowl cleaner. But also by externalising harm/pollution/costs, monopolising, reducing taxes, etc. All of which harm you, personally.

So, sure, worry about governments. But worry more about (big) companies. Read more history.

Mistral OCR 4 29 days ago

All AI companies are working on models with specialisms. Which are really good at one task.

Mistral is just a bit more forward about this. I guess because they don't need/want to "wow" an audience with generalist user-facing tools (chat) that seem to be experts in everything (but in reality quite often will be a lot of such specialist models chained together).

Here, what you want, is really just a few python scripts away. Voxtral to turn your spoken prompt into text, piped into mistral large 3 with extra system prompts that creates a prompt for ocr and paths to files. It could do this in a loop to actually find those files. which you throw at ocr3, is pased back to misteal large 3 to interpret and turn into decisions.

This is common. It's rather uncommon, really, to build something like this using only one model for everything.

I'd imagine the (aggressive) caching of the favicon by browsers makes it a challenge, but you could generate the favicon dynamically, then have JS extract the sequentially. Basically streaming arbitraily large content to a webpage via favicons. Via blocks of 239 bytes.

It may be a fun, novel way to proxy webpages that are otherwise blocked. Though, i guess, the service rendering the favicons can just as easily be blocked then.

An SVG can embed raster images: base64 encoded bytes.

So you could layer this experiment: favicon is svg, that contains encoded raster, whose bytes are encoded html.

At the very least it would make a mindboggling CTF step.

Same here.

I tried content-types, user-agent, but no luck. I'm not sure what the user-agent of `req` is, but the default `node-fetch/1.0` does make the response json. They are a 307, but the result is a png.

I presume the original payload may have contained information that the hackers want to keep from prying eyes. Esp. now that it landed on HN, it makes sense to take it offline and replace with an actual png to avoid people finding information in it that may harm their future hacks or so?

Probably my reference is tainting this then. I come from Perl, PHP (both of which, at that time, had nonexisting or terrible package managing - CPAN, PEAR), then Ruby, with gems and now Python, Typescript/Javascript. Starting with Ruby, I've developed in communities that heavily depend on deep dependency trees.

For me, the discipline of shallow dependencies, no-dependencies etc, was new when I came to rust.

Sure, if I pull in something like Rocket, it comes with dependencies, that have dependencies that have dependencies. But Rocket is one of the more extreme examples I know of, and even that is nowhere near the depth of a tree that common (not extreme) npm frameworks/libraries come with. Before yarn I sometimes had node-module trees that went over 50 levels deep.

The Python community doesn't have this extreme deep dependencies, but in Python it is far more common to "from foo import bar" in both libraries and in applications than to write a few hundred lines of code yourself. The Django and "lean" flask, or "simple" cli apps I worked with and on, quite commonly have hundreds of dependencies (many of which are dependencies of dependencies etc).

Within that context, rust community is far more conservative. Many of the dependencies that I use have one, maybe two of their own deps. Many none. And it's more common - IME - to see libraries that have just one level of deps - the libs a lib depends on, itself won't have deps.

Though, I guess, C or even C++ community, lacking OOTB, common and easy dependency management like cargo, will be far more conservative even.

Rather than being "flat out wrong", I'd say it very much depends (pun intended) on where you come from and compare rust with.

That's the default. One can also statically bind that libc.so.6 quite easily. Though that's not the default.

edit: Ironically, that makes shipping the binary a tad harder, since this "linux" version won't "magically" run on about every Linux, or mac version on mac, etc. I guess that's why its not the default, though that's just me guessing.

"Written in Rust" signals some common attributes.

Fast, Safe, Lightweight, Statically linked (plop a precompiled binary in ~/.local/bin and run it), few/shallow dependencies, senior developers, "Done" software.

Now, certainly no guarantees, enough counter-examples, I know. And attributes that one can get with anything from PHP via Javascript to Lisp as well. Some attributes have stronger correlation than others too.

But, in general, "rust" has a (much) higher chance of meeting these attributes. I care about those attributes above anything else.

Rust even doesnt do the static binary file by default.

Huh? It does. Only libc is dynamically linked, by default, which --iirc-- all programs will commonly need anyway. All the rest is statically linked.

In fact, it takes some hoop-jumping to build dynamically linked binaries with cargo.

I know of https://usbguard.github.io/

But I remember that on Linux changing some /etc/udev file helped me with some naggy bug long ago. I worked temporary in an office with several wonky USB keyboards. Whenever someone disconnected their tablet or laptop from their KB (ie shut the lid), my linux would pick it up and suddenly connect to this KB. A little googling and some trial-error and I had my linux set-up that it would only connect to whitelisted USB devices.

Which, months later, caused me insane headaches when I could not find why a new USB microphone wasn't working, despite it being advertised as "works on linux"....

They can fix it. They have certainly figured out how. But their "killer feature" is not that you don't receive spam, it's that the mail you send isn't flagged as spam by their fellow oligopolists.

We're now at the place where it's virtually impossible to run your own mailserver and have the mail delivered, consistently at Gmail and Outlook/Live/Hotmail. At least not without hours a month tuning, re-configuring, monitoring etc.

Basically, Gmail, Apple Mail, Microsoft, Yahoo (and to lesser extent, Fast-email, proton, or one of the handfull of dedicated email providers) have cemented an oligopoly. You must invest serious infrastructure, time and effort, or else your mail will be /dev/nulled (at random, often).

This "anti-spam" works, reasonably well. Because Gmail can now trust that Microsoft has measures in place to disencourage new accounts from sending large amounts of mails - and vice versa. Obviously Gmail can trust other Gmail accounts. And so they have a win-win-win.

win: No need for heavy, resource-intensive spam-training or scanning for the bulk of incoming mail - if its from a fellow BigTech, let it through. Win: an almost impossible high barrier to entry for any serious competitors. Win: Lock in, because anyone wishing to move will see their email not reach the inboxes of users at other Big Tech - aka the vast majority of inboxes.

Have you considered collecting all the literals into domains, but ship them by default?

I could, for example, imagine using roto in some of my current work on svg and visuals generation. In which case I'd be greatly helped with literals like "colors", "vec2", "angle" etc. I'd imagine that as long as other literals which I don't need, like an IP address, aren't in the way, it's still greatly beneficial to have a large lib to pick and choose from, around.