HN user

mbreese

14,533 karma

username at gmail.

Posts52
Comments4,188
View on HN
www.osnews.com 1mo ago

Open source project contains hidden instruction for "AI" agents: delete my code

mbreese
17pts3
bsky.app 5mo ago

Godot is drowning in AI slop pull requests

mbreese
10pts4
www.anthropic.com 9mo ago

Claude for Life Sciences

mbreese
1pts0
www.theregister.com 1y ago

Linus Torvalds hints Bcachefs may get dropped from the Linux kernel

mbreese
6pts0
ohadravid.github.io 1y ago

BTrees, Inverted Indices, and a Model for Full Text Search

mbreese
4pts0
kubernetes.io 1y ago

Introducing JobSet

mbreese
2pts0
bret.dk 1y ago

The Raspberry Pi CM5 Is Weeks Away?

mbreese
3pts2
www.percona.com 1y ago

How Can MySQL Catch Up with PostgreSQL's Momentum?

mbreese
3pts0
www.tomshardware.com 1y ago

Raspberry Pi RP2350 powered PyDOS handheld in a BlackBerry form factor

mbreese
9pts0
github.com 2y ago

SQL-Only Webapp Builder

mbreese
3pts1
www.cnbc.com 2y ago

Computing firm Raspberry Pi pops 31% in rare London market debut

mbreese
22pts3
www.pdl.cmu.edu 2y ago

DeltaFS: Rethinking Filesystems for the Exascale Era

mbreese
1pts0
seaweedfs.github.io 2y ago

SeaweedFS is a simple and highly scalable distributed file system

mbreese
3pts0
hackaday.com 2y ago

Building a better keyboard and mouse switch

mbreese
2pts1
www.githubstatus.com 3y ago

GitHub Status Incident (2023-06-29)

mbreese
14pts4
www.creativebloq.com 3y ago

Playful British Airways ads read travellers' minds

mbreese
1pts0
www.axios.com 3y ago

Elon Musk completes Twitter takeover and fires top executives

mbreese
10pts0
mymodernmet.com 3y ago

Letter with Hand-Drawn Map Instead of Address Arrives at Destination in Iceland

mbreese
3pts0
www.infoworld.com 3y ago

RStudio changes name to Posit, expands focus to include Python and VS Code

mbreese
5pts0
support.starlink.com 4y ago

Starlink Adds Portability Option

mbreese
1pts0
itwire.com 4y ago

Debian developer demoted, quits after two decades with project

mbreese
56pts30
arstechnica.com 4y ago

Developer sabotages his own apps, then claims Aaron Swartz was murdered

mbreese
6pts3
www.theverge.com 4y ago

Apple won’t let Epic bring Fortnite back to South Korea’s App Store

mbreese
3pts1
slate.com 4y ago

History of Dean Kamen’s Segway

mbreese
2pts0
www.extremetech.com 5y ago

Intel Discontinues Lakefield, Its First x86 Hybrid CPU

mbreese
23pts0
www.tomsguide.com 5y ago

Tesla just announced a brand new car - a $25K hatchback

mbreese
18pts16
www.forbes.com 5y ago

There Were No Borders in the Middle Ages

mbreese
4pts0
www.docker.com 5y ago

Docker Desktop for Mac (Apple Silicon)

mbreese
247pts122
blog.jfedor.org 5y ago

Bluetooth Trackball Mark II

mbreese
409pts95
yarh.io 5y ago

Yarh.io Micro 2 Raspberry Pi 3B+ Hacker's Linux Handheld

mbreese
1pts0

This was my biggest question. Thanks for preemptively answering (and being proud!).

When I was playing with this, I immediately thought of TiddlyWikis, which I've always been fond of, but they never took for me. However, an easily editable presentation -- that's a use case I can understand.

I couldn't see how there was a CDRT when I'm just working with a single file. I was trying to figure out where the server was. So -- this is just on for everyone? Are you concerned that your Cloudflare account will be hammered with this?

The linguistic gymnastics required when talking about OpenAI vs ChatGPT and Anthropic vs Claude is difficult when you're giving talk about them. At least Google vs Gemini is a little clearer.

I mean, I get the rationale Company vs. Product, but most people know the product. As in "I used ChatGPT". But if you ask who OpenAI is, they'll have no clue.

ChatGPT is in someways nicer... because their models are GPT-5.3, GPT-5.4, etc...

But when you're trying to explain that the Anthropic models are called "Opus" or "Sonnet" or "Haiku" or "Fable", but you use them in "Claude", it gets confusing quickly.

I’m not sure you’d see that big of a difference in quality. There is quite a bit of cruft that can accumulate when you know the code will never be public.

But, you would probably see a difference of scale and architecture. Larger projects that need better organization are probably more likely to be in private codebases (Linux excluded). So you might be right about the lack of private code in LLM being an issue.

A simpler spec can be used by a simpler agent. So, maybe that's the use-case here... use by smaller/cheaper agents that run in parallel as opposed to large models running one visualization at a time.

Or at least, maybe that's the idea?

IME, Claude and ChatGPT do just fine generating ggplot models, but extensive customization can get a bit hairy.

They just need to price it more than the cost of the base stations and maintaining downlink sites. Any revenue above variable costs is worth trying to capture. The cost of keeping the constellation of satellites healthy will be borne by customers in richer regions.

Look at it this way - the satellites will be traveling over Africa no matter what. If no one in Africa subscribes, then that's a part of their orbit where they aren't earning any revenue. Assuming the rest of the costs are covered by subscribers in other regions, they can price service in Africa as low as they'd like (above hardware costs).

For what it's worth, I was specifically thinking about DSL and 2000-2010s era DOCSIS technologies. The analogy is really that when mobile phones became popular in Africa, it was possible to skip the step of having wired landlines and cable TV. 3G/4G could be used for anything those would have provided. Starlink does the same thing, but it lets people skip DSL and cable modems.

However, I actually think that the main "benefit" you get by not having an extensive wired infrastructure is the lack of an incumbent provider. Centralized incumbents can put pressure on any new technology that puts their existing infrastructure investments at risk.

Isn’t this a similar argument to how Africa adopted mobile phones significantly faster than other regions? When you don’t have an established wired infrastructure, it becomes significantly easier to jump technology generations. Especially if there’s no infrastructure needed to install.

As others mentioned, It’s a very similar situation for rural America. My dad lives in a rural setting, and for years could only get slow geostationary satellite Internet. As soon as he got Starlink, his connectivity improved dramatically. Only now that there was an established market for rural internet users in his area, are cable and fiber lines starting to get run.

barring physical limitations

I think you're missing the point. This is a physical limitation. People don't write as much as they used to. I used to be able to write for hours, taking notes, writing papers, etc... Now, I try to take notes during a meeting and my hand cramps. And as a bonus, my handwriting never was great, but now it's illegible for me.

Yes, we all "could" re-learn this skill, but how many people will? If you asked me to type a paper on a locked-down computer, I could easily. If you asked me to write a 2000 word paper by hand in an hour, I doubt it would happen.

If we expect students to take in-person tests on paper, then you should also do the rest of the classwork on paper. I am completely in the camp that we learn more when we write something. The physical act is part of the learning process. But you can't expect students to write an in-depth exam without having the practice of doing it often.

Biotech and academia have very different standards for data quality and reproducibility. Most of the biotech people I know view academic research as an interesting first draft at best.

Or Claude is better at getting people to move to more expensive membership tiers. From reading here, it seems like Claude still has a lot of users. If Claude has lower limits for their $20 plan, it stands to reason that people are paying for more expensive plans to get similar levels of usage. This assumes they aren’t reducing demand through the throttling, which is a big assumption.

I’d love to know what Anthropic’s comparable numbers look like.

If we're talking a government site, chances are you don't have the budget to be able to hire much above junior or midlevel devs. And the project manager probably has a small budget [^1] and little experience with what the web design choices really mean (and what the trade off are).

I think you'd be surprised who ends up making those decisions.

Which goes back to the original point (that's valid for any project) - keep your user in mind. If your users will be using recent-ish iOS or Android devices, use as much flair as you'd like. If your users will be using mass-market low-end devices or used devices from 4+ years ago, then maybe dial down the interface.

Knowing your user is important, no matter what level you're at.

[1] Unless we're talking about some kind of large system that's being redesigned by a consulting company on a cost-plus contract. Who knows how those decisions are made.

I'm somewhat in agreement. I like building 1:1 code for that specific agent.

Where I'm starting to question this is maintainability. When I come up with a new technique or way of doing something in my new agent, how can I update an older agent. Do I want to update the older agent?

But, I get what you're talking about w.r.t. building for the exact problem at hand. For example, I'm guessing that Apache Burr has support for a plugin-able vector RAG system (or at least it will if it doesn't now). That's great, but I want my RAG system to add documents to the context and keep them as part of an updated system prompt with some very specific tweaks that happen as part of that process. This is a bespoke way of working with an existing concept (RAG) that doesn't lend itself to using any specific framework.

In my use-case, bespoke is the way to go. But then I'm still stuck with having to make engineering choices for updating older agents. So, I see your point.

I’m not sure I agree with all of your takes either. For example, I’m not anti-AI for coding, so that immediately made me click away. I’m glad I read the comments though because I think the take of “not using your code to train AI” makes a lot more sense.

But, I wanted to say thanks for posting this and being really open in the comments. It’s hard to get so much feedback so quickly. It’s a firehouse of criticism that’s hard to deal with.

You’re handling it well.

I’m not going to defend Fisher here. It was a stupid thing for someone to do.

But unless you’re in the field, you won’t realize exactly how big ThermoFisher actually is. They are the major supplier of everything for molecular biology work. From freezers (the Thermo part) to plates and pipettes (Fisher) to enzymes and antibodies. In many ways they are like Amazon. They sell everything. Some of it from outside companies, but a good deal of sales are from in-house brands. They could use their position as a reseller to know which products sell the best and with the highest margins.

In a company of this size, it’s easy to have one group feel pressure and cheat on running the gels to confirm results. Particularly when the real results are ambiguous or dodgy. It’s not a good look, but I doubt it will put a dent in people from buying things (non-antibodies) from them.

This is the thing. Yes, the marketing material is bad. But, no one in lab trusts an antibody just because of where you bought it. A new antibody always gets tested and validated before use.

That is to say, this looks bad for Thermo Fisher. But, that’s as far as the damage should go.

I have to agree. I’m interested in the project, so congrats. It’s something I might really like using.

But the one thing I expected to see in the Readme was an example of: takes this tool run output: XXXXXX and converts it to: XX for a savings of 40% of tokens.

This looks like a nice (and useful) project, so thanks for sharing!

I think it’s a question of cost/benefit.

For the researcher, it can take a lot of extra time and effort (and skill) that they might not have. A unoptimized job that takes four days to run is still faster than taking a week to optimize the code to run in 1 day.

For the researcher, the main limit is time. In many places the cost of the HPC hardware isn’t passed onto them, so their main pressure is time. And running code is generally faster than optimizing code.

(Unless you’re running a week long analysis thousands of times)

Thinking of this as an allocation program for the application to manage is an interesting approach. But the program will need to be able to model their resource requirements from start to end, and know about how long each step will take. This sounds like a variant of the halting problem, but instead of predicting when a program will end, it’s predicting when it will need more resources.

Yes it is common, just not always to the scale of a Veritasium video. Usually it’s just the press office for a university putting out a press release or a summary article in Scientific American.

But in the case where the story is interesting to a larger audience, having a push behind a story across non-academic media is not unheard of. If you can get some media coverage of an academic topic, it can be very beneficial to the researchers’ careers. One goal for a researcher is to bring notoriety to their research, to their institution, and to the field in general. This is the main motivation I see.

The authors may have pushed the arxiv paper out earlier due to the timing of the release of the video.

I don't think I've ever installed a printer's app on my Mac. I have a reflex against it. The Mac has decent support for printers and scanners built-in, so did you need an App at all?

That's one thing that bugs me about hardware companies -- they all want you to install their app to monitor/configure/bedazzle their hardware. But really, it probably is just going to work anyway. And a printer or a mouse shouldn't need much.

At least on a Mac. On Windows, sometimes the vendor apps are helpful, but usually it's still not absolutely needed (except maybe for GPUs...)

Many HPC jobs aren’t simulations that are CPU bound. In fact, most of the jobs on the clusters I’ve used have been single-node jobs (so technically HTC, but that term is rarely used).

I do genomics work and my jobs tend to be bursty. They may use a lot of CPU initially, but the second half of the job is writing results. This takes only one core, but still the max amount of memory. Or, I can have jobs that are CPU light, but need the max amount of memory for only a fraction of their wall-time.

Here is an example for you. Let’s say I’m processing a genome sequencing experiment. This requires about 8 different steps between preprocessing the data, alignment, post filtering, QC stat collection, etc. These are large input files, so my jobs end up being IO bound. If I were reading and writing at each step, it would add days of time to the pipeline. Instead, what we do is read the data once, and pipe the data from program to program. But each program has different CPU and memory requirements. We need to reserve the $MAX requirements for each. As the data moves through the pipeline though, we eventually end up with max utilization for only a portion of the walltime. If I optimize for efficient walltime, I leave CPUs and memory idle for a large portion of the job.

People also tend to like to manage fewer jobs. So instead of splitting a job into multiple dependent submissions that are tailored to each program, people will write a bash script that runs, but is not efficient.

Many times, these patterns are difficult to predict and you can’t submit a job that says I need 20 cores for an hour, but only 2 for the last two hours. It’s difficult to balance utilization vs wall-time. No one likes waiting, but there is usually little incentive to have high utilization rates. And sometimes the balance is total walltime. Sometimes it’s execution complexity.

This is the problem this group is trying to solve - dynamically adapting the scheduler to know when a job isn’t going to use its full allocation of resources. I’m not sure there is a good way to do it. HPC users are concerned only with getting their jobs done fast. HPC admins want to see resources used efficiently. This is a classic pipelining problem: do you optimize for individual task time or overall system throughput?

I think the only way to really do this well is to make HPC jobs a market system where resources cost money to the users. When money is involved, people are incentivized to optimize their workloads. But that’s rarely the case for large HPC clusters and I’d personally hate it if I had to deal with a HPC processing budget.

In lieu of this, a common way to handle this lack of efficiency is to do “fair share” scheduling. This means that a users prior work load is taken into account when prioritizing their queue position. So, if I did a lot of work last week, jobs for a user that didn’t run jobs last week would get a priority boost over me. This doesn’t address the utilization efficiency directly, but it does make access to the cluster seem more “fair”.

Nvidia RTX Spark 2 months ago

I'm not sure that's such a bad thing. It's not going to challenge the Apple M5, but if you're specifically looking for something in the "not-Mac" market, having a laptop-sized version of the DGX is probably going to be pretty successful.