HN user

jballanc

9,395 karma

I'm just a guy with some ideas...programming lets me do something with them.

Get in touch! jballanc@gmail.com https://github.com/jballanc

[ my public key: https://keybase.io/jballanc; my proof: https://keybase.io/jballanc/sigs/pxb6FPsVF9hoQMyb2ap0euWoNG3YxwN2g4VY94RRxX8 ]

Posts114
Comments1,418
View on HN
zenodo.org 1mo ago

Adaptive Low-Rank Transformer with Dynamic Expert Routing for Continual Learning

jballanc
3pts0
zenodo.org 2mo ago

Show HN: RVW – A transformer model capable of online continual learning

jballanc
1pts0
news.ycombinator.com 3mo ago

Ask HN: What would you do with an AI model capable of continuous learning?

jballanc
4pts6
news.ycombinator.com 3mo ago

Ask HN: Why don't frontier AI model providers continuously improve their models?

jballanc
1pts1
www.manhattanmetric.com 4mo ago

The Singularity Is Coming

jballanc
2pts0
news.ycombinator.com 4mo ago

Ask HN: Where should an independent researcher publish work on ML?

jballanc
2pts2
vibe-lang.org 4mo ago

Show HN: Vibe – a language for humans and AI to reason about programs together

jballanc
2pts0
www.manhattanmetric.com 4mo ago

LLMs – How did they get so good?

jballanc
1pts0
www.manhattanmetric.com 4mo ago

LLMs – What aren't they good for?

jballanc
5pts1
www.manhattanmetric.com 4mo ago

ChatGPT Told Me to Go Work for Anthropic

jballanc
4pts1
news.ycombinator.com 2y ago

Ask HN: Has anyone else found it harder to review code recently?

jballanc
28pts40
gaiwan.co 4y ago

Still Beating the Averages

jballanc
2pts0
mjtsai.com 8y ago

Apple Officially Discontinues Airport Router Line

jballanc
1pts0
medium.com 8y ago

Anti-Gun, Pro-Second Amendment

jballanc
1pts1
juliacomputing.com 9y ago

Julia Enters TIOBE Top 50 Programming Languages

jballanc
1pts0
www.hurriyetdailynews.com 10y ago

More Turkish women consider writing code as career

jballanc
1pts0
arxiv.org 10y ago

Coinami: A Cryptocurrency with DNA Sequence Alignment as Proof-Of-work

jballanc
2pts0
www.kickstarter.com 10y ago

Prove if your luggage was mishandled by the airlines

jballanc
1pts0
blog.fikesfarm.com 10y ago

Planck is a ClojureScript REPL and script execution environment

jballanc
86pts10
dev.clojure.org 11y ago

Reader Conditionals in Clojure

jballanc
91pts10
www.psychologytoday.com 11y ago

How a Generation Was Misled About Natural Selection

jballanc
2pts0
www.rubymotion.com 11y ago

RubyMotion Success Story: Freckle

jballanc
17pts0
www.rubymotion.com 11y ago

New Pricing Plans for RubyMotion

jballanc
1pts0
blog.openmicroscopy.org 11y ago

On Being a Partner or “Send Us the Data”

jballanc
4pts0
blog.openmicroscopy.org 11y ago

A frustratingly incomplete, always obsolete, amazingly effective solution

jballanc
1pts0
blog.openmicroscopy.org 11y ago

The Joy of File Formats

jballanc
3pts1
unicornfree.com 11y ago

Startup Winter is Coming

jballanc
15pts6
futurestack.com 12y ago

A Conference Stole My Identity

jballanc
237pts142
github.com 12y ago

FizzBuzz implemented with Clojure's core.logic

jballanc
2pts0
medium.com 12y ago

Why am I Angry with the Video Game Industry in Turkey?

jballanc
4pts0

I've been working on RVW, my adaptation of the standard transformer model that is capable of online continual learning without catastrophic forgetting. I finally published the first pre-print of my early experiments: https://doi.org/10.5281/zenodo.20064617

Now I'm working on expanding the work into more parameters and improving performance. I just finished an extremely harsh test of a Nemotron-flavored RVW that consisted of stretches of a random assortment of domains interspersed with long runs of single domains. Across all of it the model didn't forget (and actually improved on some of the more challenging domains). PPL on SmolTalk is still in the ~18 range, which I'd like to get lower, but this is all with only 4B params.

Currently, I'm training a Llama 3.2-flavored RVW with only about 2B params to see how that turns out. Depending on results of that, I may take it to Gemma 4 next.

NetHack 5.0.0 3 months ago

IIRC, there was always a way to filter out certain messages (or that may be an alt.org customization, but it's been a part of my config file for a while now).

NetHack 5.0.0 3 months ago

For real! Valkyrie is the perfect "just bash things while only half paying attention" class. Great for when I'm playing to unwind (as opposed to playing as a challenge to myself).

At least there's still Samurai.

I think Douglas Adams had one of the best quotes regarding observing infinity:

"Infinity itself looks flat and uninteresting. Looking up into the night sky is looking into infinity – distance is incomprehensible and therefore meaningless."

It's been a while since I worked at Apple, but back in the day the entire OS X Server team made extensive use of kerberized NFS shares for moving around large files...

...the last version of Server shipped in 2021 (and the last real version shipped almost a decade before that).

My first job after finishing my undergrad degree was performing quality analysis on corn starch. As a condition of employment, I had to sign a paper saying anything I invented related to corn was property of my employer.

It's been more than a few years since I worked at Apple, but they were always unique in the tech space in that their retail division dwarfed headcount. If I recall correctly all of OS X Lion was produced by around 3,000 engineers (and probably less, since I think that count included iLife and iWork).

I've been working on an ML model capable of robust continuous learning, resistant to catastrophic forgetting without relying on replay, an external memory system, or unbounded parameter growth. Last week I confirmed the first non-toy, 580M parameter version soundly beat LoRA, EWC, and full fine tuning. This week I'm scaling up to 4.4B parameters...

Based on what you've already mentioned, there's a good chance you're familiar, but on the off chance you're not: "Funkungfusion" (or, really, anything off the Ninja Tune label) might be right up your alley.

Arm AGI CPU 4 months ago

Eh, I'm not so sure it'll be that big a deal. The whole supply chain is so twisted and tangled all the way up and down. Shuffling out one piece doesn't seem like it will, on its own, be so major. Samsung made the chips for the iPhone, then made their own phone, then Apple designed their own chips made by TSMC, now Apple is exploring the possibility of having Samsung make those chips again.

Also, it takes a willful ignorance of history for ARM to claim this is the first time they've manufactured hardware. I mean, maaaaybe, teeeeechnically that's true, but ARM was the Acorn RISC Machine, and Acorn was in the hardware business...at least as much as Apple was for the first iPhone.

I exited academia for industry 15 years ago, and since then I haven't had nearly as much time to read review papers as I would like. For that reason, my view may be a bit outdated, but one thing I remember finding incredibly useful about review papers is that they provided a venue for speculation.

In the typical "experimental report" sort of paper, the focus is typically narrowed to a knifes edge around the hypothesis, the methods, the results, and analysis. Yes, there is the "Introduction" and a "Discussion", but increasingly I saw "Introductions" become a venue to do citation bartering (I'll cite your paper in the intro to my next paper if you cite that paper in the intro to your next paper) and "Discussion" turn into a place to float your next grant proposal before formal scoring.

Review papers, on the other hand, were more open to speculation. I remember reading a number that were framed as "here's what has been reported, here's what that likely means...and here's where I think the field could push forward in meaningful ways". Since the veracity of a review is generally judged on how well it covers and summarizes what's already been reported, and since no one is getting their next grant from a review, there's more space for the author to bring in their own thoughts and opinions.

I agree that LLMs have largely removed the need for review papers as a reference for the current state of a field...but I'll miss the forward-looking speculation.

Science is staring down the barrel of a looming crisis that looks like an echo chamber of epic proportions, and the only way out is to figure out how to motivate reporting negative results and sharing speculative outsider thinking.

When I was a young kid, my mother was a “stay at home mom”, which meant that she babysat the kids of 5 or 6 of the other families in our neighborhood where both parents worked. For me, it was a wonderful experience growing up having a ready-made group of close friends and my mother close at hand. It did mean that my mother effectively sacrificed her career (though she eventually went to work for my father as his office manager and was instrumental to his success), but I’m certain she was not charging $20k/yr/kid (or whatever the equivalent in 1980s dollars would be).

What Americans seem to only just now be waking up to is that lack of work/life balance, lack of family leave accommodations, and loss of community has a very real, very tangible dollar amount cost. I’m very, very tired of the knee-jerk response to every “socialist” proposal being, “yeah, that’s great, but how are you going to pay for it?”

How are you going to pay for not having family leave? How are you going to pay for not having universal healthcare? How are you going to pay for not having tuition-free college for all? These choices have a cost, and Americans are paying that cost every day!

Facebook is cooked 5 months ago

Reporting blatant criminal violations is not the same thing as moderating otherwise-protected speech that could be construed as misleading, offensive, or objectionable in some other way.

Facebook is cooked 5 months ago

The problem with this is that section 230 was specifically created to promote editorializing. Before section 230, online platforms were loath to engage in any moderation because they feared that a hint of moderation would jump them over into the realm of "publisher" where they could be held liable for the veracity of the content they published and, given the choice between no moderation at all or full editorial responsibility, many of the early internet platforms would have chosen no moderation (as full editorial responsibility would have been cost prohibitive).

In other words, that filter that keeps Nazis, child predators, doxing, etc. off your favorite platform only exists because of section 230.

Now, one could argue that the biggest platforms (Meta, Youtube, etc.) can, at this point, afford the cost of full editorial responsibility, but repealing section 230 under this logic only serves to put up a barrier to entry to any smaller competitor that might dislodge these platforms from their high, and lucrative, perch. I used to believe that the better fix would be to amend section 230 to shield filtering/removal, but not selective promotion, but TikTok has shown (rather cleverly) that selective filtering/removal can be just as effective as selective promotion of content.

More like it's time for the pendulum to swing back...

We had very decentralized "internet" with BBSes, AOL, Prodigy, etc.

Then we centralized on AOL (ask anyone over 40 if they remember "AOL Keyword: ACME" plastered all over roadside billboards).

Then we revolted and decentralized across MySpace, Digg, Facebook, Reddit, etc.

Then we centralized on Facebook.

We are in the midst of a second decentralization...

...from an information consumer's perspective. From an internet infrastructure perspective, the trend has been consistently toward more decentralization. Initially, even after everyone moved away from AOL as their sole information source online, they were still accessing all the other sites over their AOL dial-up connection. Eventually, competitors arrived and, since AOL no longer had a monopoly on content, they lost their grip on the infrastructure monopoly.

Later, moving up the stack, the re-centralization around Facebook (and Google) allowed those sources to centralize power in identity management. Today, though, people increasingly only authenticate to Facebook or Google in order to authenticate to some 3rd party site. Eventually, competitors for auth will arrive (or already have ahem passkeys coughcough) and, as no one goes to Facebook anymore anyway, they'll lose grip on identity management.

It's an ebb and flow, but the fundamental capability for decentralization has existed in the technology behind the internet from the beginning. Adoption and acclimatization, however, is a much slower process.

Yeah, "foolish" maybe wasn't the right word. All metaphors fall short in some way (hence why they're metaphors). I just, knowing something of the history of that part of the world, like to use the opportunity to share the knowledge that, despite the appearances of a chaotic, random aggregation of humans, Bazaars often had a significant structure under the surface (perhaps another lesson about open source to be had there).

It's a reference to Eric S. Raymond's famous article "The Cathedral and the Bazaar", where he compares the rather top-down, leader driven culture of Unix development to the free-for-all style of Linux.

Of course, I always like to point out the foolishness of this metaphor: Bazaars in the Near East were usually run in a fairly regimented fashion by merchant guilds and their elected or appointed leaders.

I can tell you the same thing I was told when I started my program: no thesis represents more than 1 year's worth of work. The reason it takes most Ph.D.s 5-10 years (8 in my case) to graduate is that you have to fail, and fail, and fail again for 4-9 years before you find your thesis project.

In my case, I started on two exploratory gene knockout "fishing expeditions", both of which didn't turn up anything interesting after a year. Then I crystalized a protein and submitted it to X-ray diffraction, but the results were not good enough for a "high quality" structure, and besides the structure we did find was not particularly interesting. Then I switched to working on NMR structures, but ended up switching universities (politics...there's going to be lots of politics) before that went anywhere.

At my new university I switched to structure modeling and worked on a project my advisor suggested for about a year to optimize a modeling routine, but even the optimized version didn't turn up anything interesting. Finally, I landed on a very intriguing problem that could have had far reaching implications. I worked hard at it for almost a year, only to realize that even state-of-the-art modeling was at least a decade away from being able to begin to address the problem I needed to solve. Finally, I returned to a question that a professor had asked me in my first year of graduate school, half jokingly, assuming there was no way to answer the question. For about a year I worked hard at it, finally arrived at a very interesting answer, and graduated.

I'm no expert on Tor, but IIRC the story is precisely that spies operating from hostile territory would have a red target painted on them from using encrypted communications...unless a whole lot of people in that hostile territory were also using encrypted communications. This is why Tor was released open source and wide adoption was encouraged.

Nice write-up, but no discussion of Ruby's Range class is complete without at least a mention of the venerable flip-flop!

Can you predict the output of the following?

    (1..20).each do |i|
      puts i if i.odd?..i.prime?
    end

Without having read in-depth either original paper, it seems like the issue here is much simpler than reproduction (though reproduction is the gold standard as is totally under-appreciated these days).

Rather, it seems the authors made a much simpler mistake: hypotheses can only be refuted by evidence, not confirmed. So, in this case, if the hypothesis is "judges act more harshly when hungry", what they should have been doing is looking for evidence disproving that statement. Instead, they seem to have presented a correlation and a suggestion, which is not the same thing as a scientific finding.