HN user

bnjemian

401 karma

AI/quantum scientist, 0-1 software engineer.

Posts8
Comments69
View on HN

It’s a huge problem, but I’d caution against this absolutism — there may well be structure that can be created around and between LLMs and their outputs to enable the necessary segregation.

As a loose comparison, hardware bit errors happen probabilistically, yet they’re so rare that we can effectively ignore them in day-to-day use assuming no specialized application (e.g. defense, space, critical infrastructure).

LLMs aren’t there yet, but it’s entirely plausible that structures may can be developed to solve the problem, and those structures aren’t known or commonly conceived of in the present.

Location: Portland, OR

Remote: Remote/Hybrid/On-Site

Willing to relocate: California, Washington, Canada, EU, or UK (no visa required)

Languages: Python, Rust, TypeScript, JavaScript, HTML/CSS, SQL

Technologies: PyTorch, JAX, Tensorflow, CUDA, LangChain, LangGraph, NumPy, SciPy, Polars, Pandas, Numba, Cirq, QISkit, Pennylane, Node.js, Deno, Express, Hono, PostgreSQL, PL/pgSQL, Redis, AWS, Azure, Docker, CI/CD — plus whatever you throw at me; I learn fast.

Resume: https://drive.google.com/file/d/1kdtM2vPpLmTu3HI51Guetj82CPe...

LinkedIn: https://linkedin.com/in/benjamin-cordier-phd

Background: Staff-level full-stack engineer (13 YoE), both product and research. PhD in quantum computing/machine learning, applications in computational biology. MS in bioinformatics, applications to cancer genomics. I’m also the founding engineer behind June Dating (0-to-1). Our founding team bootstrapped the company to profitability, we’re currently on a product development hiatus. For my day job, I work in Rust and am building a data watermarking system for a flagship NIH consortium centered on ethical AI and FAIR data, currently around 160M files watermarked (1.9PB of data).

Open to: Founding engineer, staff/senior-staff engineer, and engineering manager roles. Industry/domain agnostic. I do especially well in high ownership and high impact roles.

Okay sure, but what happens when a high CVE is discovered that requires immediate patching – does that get around the Upload Queue? If so, it's possible one could opportunistically co-author the patch and shuttle in a vulnerability, circumventing the Upload Queue.

If you instead decide that the Upload Queue can't be circumvented, now you're increasing the duration a patch for a CVE is visible. Even if the CVE disclosure is not made public, the patch sitting in the Upload Queue makes it far more discoverable.

Best as I can tell, neither one of these fairly obvious issues are covered in this blog post, but they clearly need to be addressed for Upload Queues to be a good alternative.

--

Separately, at least with NPM, you can define a cooldown in your global .npmrc, so the argument that cooldowns need to be implemented per project is, for at least one (very) common package manger, patently untrue.

# Wait 7 days before installing > npm config set min-release-age 7

I just experienced this same issue with Gemini. I pasted a text message thread into Gemini (Pro, Thinking, Flash – all are affected) and it was misattributing dialogue. It said Alice said x, which Bob had said; it said Bob said y, which Alice had said. This was a two person dialogue and clearly marked with:

Alice: x Bob: y Alice : z ...

While the analysis was mostly coherent with the exception of said misattributions, I filed away the mental note that this misattribution error happened frequently in these type of exchanges.

It's funny because the author notes a prior attempt to uncover Satoshi's identity and giving up because an implied lack of technical depth.

I guess this time they were undaunted. Perhaps they received an AI assist and felt validated by AI sycophancy.

Much of the technical evidence cited is weak (e.g. strong knowledge of public-key cryptography, both used C++, etc.). Still, the (somewhat lazy) forensic linguistics is interesting.

This completely ignores that: 1. Russia was the aggressor in Ukraine, 2. Putin has made clear his desire to pursue expansionist goals through military action targeting prior members of the Soviet Union, 3. Putin regular threatens nuclear war with Ukraine, 4. Russia has shown outward hostility towards Western democracies and sought to manipulate elections with information warfare to reach their goals (most notably, 2016 US Election and Brexit), 5. Russian regularly cuts cables connecting countries, and 6. Though completely unrelated, Putin has a history of assassinating political opponents. That's wolfish behavior if I've ever seen it.

Need to look into how this turned out – I've sent letters to Merkley and Wyden over the years about privacy concerns relating to facial recognition and similarly invasive technologies. We need more regulation in this space.

That said, the TSA is in some respects the lesser concern. Don't get me wrong, the TSA not having free rein with facial and biometric technologies is a good thing. But when companies like Clearview AI (https://www.clearview.ai) sell their facial recognition technologies to local police departments – technologies that were built on illegally obtained data and have a history of substantial racial bias – we have bigger issues. It's opaque, unregulated, invites a wellspring of social injustice, and doesn't past muster under any ELSI framework.

Government regulating government is important. But we, as a society, need to stop giving private companies like Clearview AI a pass on harmful, exploitative behavior – especially when they're run by founders like Hoan Ton-That who offer post-hoc rationalizations that amount to (and I'm paraphrasing here) 'Well, if we hadn't done it, someone else would have, so why not us?'

We need a bigger bill that enshrines and elevates privacy for the modern world.

I once read that some people who are blind from an early age, as they get older, start to click their tongue, but often those around them (parents, siblings, etc.) will discourage them. Thing is, that clicking can actually be used to develop a type of vision that operates similarly to echo location in cetaceans (whales, dolphins, etc.) – it comes about because the child realizes that if they make a sharp sound, they can begin to orient themselves with the reflections of the sound waves. After all, vision is in the brain; the eyes are just the sensors. Point being, if your son starts making clicking sounds with his tongue, you likely won't want to discourage that. And on the flip, teaching him to click may provide a means of developing his vision in an alternative way.

Edit: Here's a Pubmed article on a study where blind and sighted people were trained to echolocate: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8171922/

I don't know if this is true – that pupil sizes vary meaningfully between races and folks from Africa and Aboriginal populations in Australia have smaller pupils – but it may make sense. Those are both relatively sunny places; Northern latitudes are less so. Greater dilation (or dynamic range around the dilation), more light, possibly improving certain aspects of vision in low light. Of course, the inverse may also hold – less ability for pupils to constrict in very sunny places would be problematic too. And yet, I say this knowing that hypotheses derived from first principles and uninformed of biological context tend to be very low mileage in the biological sciences. Biology is rarely so simple.

Here's the thing you may be missing: The complete diversity of human phenotypes (including what is socially discussed as 'race') is almost entirely present on the African continent. If you believe in evolution (and I'm assuming pretty much everyone here does), that makes a whole lot of sense – humans migrated out of Africa millenia ago and, as they moved to different environments, preferential selection for certain traits that already existed within the migrating population(s) occurred. There may be some traits that are beneficial and passed on due to spontaneous mutations post-migration, but they are relatively few and typically present in superficial features (e.g. eye color, hair color).

Mark Z. Jacobson! Haven't heard that name in a few years.

Don't know him personally, but here's a tangent for the interested: In the first year of my PhD, I read several of his papers from the 90's on the GATOR family of climate models. At the time, I was interested in a potential intersection with my field. One thing that struck me was the absolutely exquisite attention to detail in one of his papers made to model the perspiration of water vapor from leaf stomata in forests (don't have the paper handy but can find if anyone's interested). It was really quite impressive.

Anyway, just thought I'd share the anecdote :)

Been using Hono for a few months and have really enjoyed it. For me, it's been the perfect minimal HTTP/router functionality needed to structure an API that lives within a single edge function that's deployed to Supabase and interacts with the Postgres DB therein. Great project – simple, intuitive, fast, lightweight.

Yes, that's somewhat true, but in practice we have subtypes. As a counterfactual to that assertion, if it were meaningfully a billion different things, then we would need a billion highly precise treatments. Yet, we've managed to do decently with relatively few.

In principle, yes, in practice, no; real-world mutations are (more often than not) non-random and their frequencies can be affected by a variety of factors. For example, the location of the mutated gene or region within the bundled chromatin structure inside the cell nucleus (this structure is highly conserved into what are known as topologically associated domains, or TADs), or the interaction between a region of DNA and cellular machinery that increases the likelihood of some mutation. There are tons of examples.

In practice, we've now molecularly characterized most well-studied cancers and know that they tend to have the same mutations. For example, certain DNMT3A mutations are very common in AML and the BCR-ABL fusion protein in CML (and results from an interaction between chromosomes 9 and 22 that produces the mutant 'Philadelphia chromosome'). There are even a wide range of cancers that share similar patterns of mutations and fall under the umbrella of 'RAS-opathies', which all exhibit some kind of mutation in a subset of genes on a specific pathway related to cell differentiation and growth. Examples include certain subtypes of colon cancer, lung cancer, melanoma, among many others.

More generally, when a cancer is subtyped, that subtyping is always done with respect to some quantifiable biological trait or clinical endpoint and – as you've hinted – that subtyping is commonly a statistical assessment. Each cancer is unique and, even within an individual cancer, we have clonal subpopulations – groups of cells with differing mutations, characteristics, and behaviors. That's one of the reasons treating cancer can be so challenging; even if we eliminate one clonal population entirely, another resistant group may take its place. The implication is that cancers that emerge with post-treatment relapse are often 1. more or completely resistant to the original therapy, and 2. exhibit different behaviors and resistance, often to the detriment of the patient's outcome.

I came to see if this had been posted given I hadn't seen it. Surprised this didn't get more attention, it's a very impressive result – and not just because it involved arrays of entangled qubits flying around each other. I watched Mikhail Lukin present this result at a Berkeley EECS seminar earlier this week. It's very compelling work and, as someone with one foot in the QIS field, I've been thinking about it quite a bit. A few observations for any future finders of this post:

- Neutral atoms, while always compelling, are now strongly in the running for state-of-the-art qubit technologies and may well have a durable superiority to other qubits (e.g. ion traps, superconducting, spin qubits, photonics). The photonics used for quantum control appear to have very powerful advantages over physical wire-based control common to spin and superconducting qubits in particular.

- The past few years, since 2019 really, have been incredibly exciting on the experimental side of the field (not that the TCS hasn't been exciting too – quite the contrary). Still, for me, this result and its timing is among the most surprising in this 4 year period. I don't I"m unique in this and I suspect that if you'd asked most folks in the QIS field on December 5th when we'd have a FTQC that can do something a classical HPC can't – even if not practical – they likely would have said somewhere between 5-15 years. Now, the path towards truly practical FTQC is clear and, with the accelerating progress, we'll likely see meaningful scientific advances due to QC technologies by the end of this decade, likely earlier.

- QC is, in many ways, a trailing technology to AI and quite exotic. While the use cases are different, the fact of the matter is that the advances in AI methods, and LLMs in particular, threaten to eat QC's lunch in many areas of scientific computing. Further, there are properties of quantum information that challenge many potential applications (e.g. no cloning theorem). In my mind, this is the greatest risk to QC technologies not gaining wide spread adoption across many STEM over the next 20 years.

- Even though QC may be practically challenged relative to AI, it is nonetheless (and likely will be for many decades to come) an incredibly verdant technological foundation for algorithm development, condensed matter physics, cosmology, etc. The quantum information paradigm is very different than the classical information; those differences provide a powerful lens to help us understand our world and universe.

Altogether, the future looks bright.

I find this fascinating. I once was considering getting my whole genome sequenced from Nebula but then I realized that their sequencing was being done by Beijing Genomics Institute (BGI). The quality of sequencing from BGI is very high, but the reality is it has some fairly established ties to the PLA (to the extent that it's considered for sanctions by lawmakers). There's been concern in the bioinformatics community for years that the CCP/PLA, through BGI, has been amassing substantial genomics data on people (most notably pregnant women). Unclear why, but – given the authoritarianism – the mind can go places. That's the thing with running any genomics-oriented organization; at minimum, pretty sordid optics can emerge if you're not careful.

I bring this up because Nebula is clearly making a play for anonymity which, in my mind, is strange. Perhaps deceitful. The thing about genetic data, especially WGS, is that it's hopelessly not anonymous. Sure, it can be protected, but given the world we live in (hacks, data brokers, etc.), it's the type of data that will likely get out into the wild. This past week at 23 and Me is a case in point. The comment from another poster here about whether they delete it or not is an important one. I suspect they don't; for a while Nebula was touting a setup where they held onto your data and you could license it for use by pharmaceutical companies or researchers and receive payment. Not sure if that's still a thing. Either way, I find the notion of "anonymity" here very dubious, even if you pay with Monero or similar, use a PO box, etc. But that's kinda the reality – these genomics companies aren't especially open about a several key facts when you get your genome sequence. Another example is that they don't make clear that when you get your genome sequenced you are effectively unmasking the genome of your extended family.

I like this idea and the design – the gamification is a good approach in my view.

A question for you: I have a relative who has been recovering from a pneumonectomy – is this a use case that you currently support or plan to support in the future (i.e. are the breathing exercises on the app similar to the ones prescribed for these patients)?

Location: Portland, OR

Remote: Yes (not required)

Willing to relocate: Yes, for the right opportunity.

Technologies:

- Languages: Python (primary), JavaScript, HTML/CSS, C/C++, R (not preferred).

- ML: PyTorch (primary), Tensorflow, Scikit-learn, Jax

- Quantum: Pennylane (primary), Qiskit, Cirq, Tensorflow Quantum

- Genomics: GATK, samtools, BWA, BioPython, Reactome

- Scientific computing: NumPy, SciPy, Pandas, Numba, CUDA, Slurm

- Visualization: D3.js (preferred), GSAP, matplotlib, ggplot2

- Web: Knockout.js, Vue, Express, Node.js, Flask, etc.

- DevOps: Git, TravisCI, Docker, etc.

- Databases: Postgres, MySQL, MongoDB, Neo4J

- Cloud: Primarily AWS, Heroku (RIP)

Résumé/CV: Upon request

Profile: I'm a current senior computational biologist, recent PhD graduate, and previous full-stack web developer. I have extensive experience in quantum machine learning, deep learning, and cancer genomics. My preference is to continue working in the quantum space, however, I'm open to exciting opportunities in the AI and biotechnology spaces, too.

I'd be very interested to see the breakdown of input energy costs. Most notable is the raw energy cost required to power the lasers and control machinery in the experiment. But then there are other costs, all of which must be amortized over time for any real-world use case to exist. I say this because the journalists in this piece imply that net gain is simply based off of the amount of energy pumped into the experiment while it operated, but the total input energy would clearly be more than that.

On the extreme end, there's the energy cost of building the machine and engineering its components. For the vast majority of these, we can probably all agree that were a fusion power plant to be built, the net gain would fully eclipse these initial inputs fairly quickly. This may sound silly, but remember that the economic context where fusion so often sits is one that centers on renewable energy and sustainability. These costs do have to be accounted for.

On the other end, there's the energy cost consumables. For example, the deuterium and tritium fuel input into the device, which need to be purified (deuterium from water, possibly tritium from the atmosphere) or otherwise isolated (from what I understand, tritium is a byproduct from fission reactors and they serve as its primary source in scientific applications). It may well be that the energy cost of acquiring these consumables is fractions to fractions of a fraction of the energy cost of running the device, effectively constituting a rounding error. But I think when we're talking about net gain, a clear definition and accounting of the input energy required to run the experiment would be useful to communicate to the public.

I hope we see disclosure of these details with all the expected caveats when the peer-reviewed article goes to print and journalists have another feeding frenzy.

Can't say I care too much about FTX's woes. Nonetheless, this article has some janky moments. The author cites an unnamed friend "who is deeply involved in [the crypto] industry" when saying CZ played a master stroke against SBF. Get a real source or at least demonstrate that your "friend" has an opinion worth 2 cents.

Maybe CZ did outplay SBF, but aside from the optics it doesn't seem anywhere that salacious. As was summed up earlier today in an NPR segment on this story "airlines don't compete on safety." In other words, it is in the interest of every exchange that the exchange industry is on the whole viewed as being comprised of trustworthy actors. Crashing FTX doesn't seem an optimal move given the already precarious state of the crypto industry these days.

Speaking of optics, the author also made a nonsensical analogy "more opaque than a dog with glaucoma" when describing the inner workings of FTX. Quite certain they meant cataracts – the clouding of the lens of the eye – not glaucoma, which is damage to the optic nerve.

Not so long ago I was a PhD student in a STEM department. With the amount of DEI and antiracist discussions had during our department meetings, it often felt that I was attending an MBA program designed for HR professionals.

Do I have anything against diversity, equity, inclusion? No.

Did I often wish the department – which was already extremely diverse in age, gender, ethnicity, race, sexual orientation, and citizenship status – would focus more on the work of doing good science, facilitating rich academic discourse, fostering cross-institutional collaborations, and driving more translation and commercial ventures? Yes. 100%, yes.

In many ways, DEI initiatives, anti-racist policies, and decolonization aspire to worthy goals. Too often, though, they feel like the product of meme-driven corporate training and consulting firms who charge wholly unequipped professionals with disparate expertise with framing their work with a funhouse mirror of their on-the-ground reality.

There has to be a better way.

I once took a golf lesson when I was a kid and, in learning about the divot removal tool, was told a similar one about the way to operate on a golf course:

"Leave it better than you found it."

I couldn't care less about golf, but this one stuck with me because it speaks to environmental stewardship. It's like a constructive riff on the mantra for when you're out in Nature: "Don't leave a trace."