HN user

panabee

4,354 karma

Founder, Hotpot.ai and HotpotBio (Hotpot.ai/bio)

X.com/panabee

Please message for free Hotpot credits. HN is an invaluable source of knowledge. I would be honored to give back.

---

Papers:

Hotpot-JOMI: Joint Outcome Model Interrogation for Reasoning Consistency

Patient.md: Framework to Organize Medical Data for AI Assistants

M.I.J.O.: Framework for Evaluating Applicability and Limitations of Biomedical Datasets with AI Assistants

Infection-Associated ME/CFS: Phenotypic Overlap, Craniofacial-to-Intracranial Sampling, and the J.O.A.N.-M.I.K.E. Frameworks

Reevaluating the Association Between Epstein-Barr Virus (EBV) and Breast Cancer in the United States

Posts421
Comments540
View on HN
academic.oup.com 1y ago

mRNA-LM: full-length integrated SLM for mRNA analysis

panabee
4pts0
huggingface.co 1y ago

DiLoCo: Distributed Low-Communication Training of Language Models

panabee
3pts0
www.educatingsilicon.com 2y ago

Revised Chinchilla scaling laws – LLM compute and token requirements

panabee
1pts0
manifestai.com 2y ago

Compute-Optimal Context Size

panabee
3pts0
arxiv.org 2y ago

An Empirical Study of Mamba-Based Language Models

panabee
43pts3
www.amacad.org 2y ago

Human Language Understanding and Reasoning (2022)

panabee
1pts0
cacm.acm.org 2y ago

MapReduce: A Flexible Data Processing Tool (2010)

panabee
3pts1
huyenchip.com 2y ago

RLHF: Reinforcement Learning from Human Feedback

panabee
1pts0
storage.googleapis.com 2y ago

Gemini 1.5 Model Family: Technical Report [pdf]

panabee
57pts3
illuminate.withgoogle.com 2y ago

Illuminate: Turn academic papers into AI-generated audio discussions

panabee
3pts1
www.cerebras.net 2y ago

Sparse Llama: 70% Smaller, 3x Faster, Full Accuracy

panabee
40pts1
www.khanacademy.org 2y ago

Mitochondria and Chloroplasts

panabee
2pts0
www.nbcnews.com 2y ago

Asian American women are getting lung cancer despite never smoking

panabee
148pts146
arxiv.org 2y ago

CatLIP: Clip Vision Accuracy with 2.7x Faster Pre-Training on Web-Scale Data

panabee
48pts4
research.google 2y ago

Patchscopes: A framework for viewing hidden representations of language models

panabee
12pts0
arxiv.org 2y ago

Understanding Diffusion Models: A Unified Perspective

panabee
2pts0
blog.adobe.com 2y ago

Fast-forward – comparing a 1980s supercomputer to the modern smartphone

panabee
2pts0
arxiv.org 2y ago

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient LMs

panabee
1pts0
www.ox.ac.uk 2y ago

Prostate cancer includes two different evotypes

panabee
152pts84
twitter.com 2y ago

Stanford researchers: 45% of GPT4 responses to medical queries hallucinate

panabee
3pts4
ar5iv.labs.arxiv.org 2y ago

Understanding Diffusion Models: A Unified Perspective

panabee
3pts0
www.sciencedirect.com 2y ago

The promises and pitfalls of specialized ribosomes (2022)

panabee
2pts0
huggingface.co 2y ago

Shallow Feed-Forward Neural Networks as Alternative to Attention in Transformers

panabee
11pts0
searchengineland.com 2y ago

New Bing attracts new Edge users – who then use Google Search

panabee
2pts0
hazyresearch.stanford.edu 2y ago

FlashFFTConv: Efficient Convolutions for Long Sequences with Tensor Cores

panabee
3pts0
twitter.com 2y ago

The core contribution of "Attention is All You Need" is logistic regressions

panabee
1pts1
med.stanford.edu 2y ago

'Anti-hunger' molecule forms after exercise, scientists discover (2022)

panabee
1pts0
leandojo.org 2y ago

LeanDojo: Theorem Proving with Retrieval-Augmented Language Models

panabee
1pts0
huggingface.co 2y ago

ConvNets Match Vision Transformers at Scale

panabee
1pts0
library.metergroup.com 2y ago

Fundamentals of Water Activity [pdf]

panabee
1pts1

Great question. The bar for proof in biomedicine is naturally high. I only shared facts because so much is unknown.

If you can find a lab exploring the question, maybe you can support them by helping to raise money for experiments.

As a fun intellectual exercise, dive into the topic and challenge yourself to think about what kind of experiments could shed more light on the subject.

For people questioning why to involve GPT and AI assistants:

GPT and AI assistants cannot be fully trusted, but they can personalize learning.

The chief challenge for the framework/handbook will be resolving how to personalize guidance into cancer research while grounding knowledge in trustworthy sources.

For instance, the framework will anchor abstract, dry biological concepts in personally meaningful tracks. Imagine someone you care about is battling lung cancer — the framework may orient learning around the molecular drivers and signaling pathways at play, or perhaps how to explore the treatment landscape while respecting established practices. If you're fortunate enough to not know someone affected by cancer, GPT can help find a personal angle.

The sheer depth of information is staggering. People devote entire careers to niche specialities, and these experts still don't know everything in their niche because our understanding of human biology and disease is constantly evolving. Adapting depth should also depend on the individual and can only be achieved via AI. Static curriculums do not maximize learning in 2026.

On second thought, I will publish something regardless of interest.

It will be an "Cancer for Engineers" framework, delivered via free, open-source Custom GPTs and Claude Skills. (Gemini gems are less reliable in our experience.)

The goal: to ease engineers into cancer via AI personalized introductory curriculums with varying time commitments to enable deeper independent investigation or fast exits if interest wanes: 4 hours, 8 hours, 12 hours.

Basically 1-3 hours per week for a month.

The reason I think some engineers may find cancer interesting, aside from the societal impact:

The human body is like a complex operating system. Cancer is a severe runtime error. Tracing root causes -- like genetic mutations, signaling errors, or immune evasion -- has many parallels to diagnosing system failures.

BTW if anyone from Kaggle/GDM is reading this, we are having issues submitting a benchmark paper for NeurIPS based on the Kaggle Benchmark.

Google models seem to get a different scheduling priority, ironically, enough and take >20 hours to complete a benchmark task that other models like Opus 4.6 finish in <1 hour -- same code path, same task. Would love help if possible since the abstract deadline is Monday (It's last minute because we didn't originally plan to submit this, but someone suggested it.)

Here are more fascinating facts about caffeine and cancer.

Caffeine affects the immune system via at least two opposing mechanisms.

Mechanism 1: A2A receptor antagonism (immunostimulatory) Tumors and damaged tissues release adenosine, which engages the A2A receptor on immune cells and signals them to stand down. Caffeine antagonizes (i.e., blocks) this receptor.

Mechanism 2: Raising intracellular cAMP (immunosuppressive) Caffeine also inhibits phosphodiesterase, the enzyme that hydrolyzes (i.e., breaks down) cAMP. cAMP accumulates inside immune cells, which acts as a "calm down" signal.

Note: both mechanisms are dose-dependent. At dietary caffeine levels, A2A antagonism likely dominates, whereas PDE inhibition is weak and mainly relevant at higher concentrations. However, the net immune effect in the tumor microenvironment remains unproven.

---

If you would like to learn more, I can outline a framework for technical folks to ease in and become more informed on cancer. Gaps abound. The more people who understand cancer, the faster we get to cures. Moreover, personalized cancer treatment is the obvious future. Knowledge acquired now may pay off later (but hopefully not needed).

If you're a wealthy person lacking a neurobiology background, how do you decide which research efforts are the most promising? Which labs do you back?

Generally, you rely on experts.

Who typically became experts by adhering to the conventional wisdom set by gatekeepers.

"Science advances one funeral at a time" feels apt.

Sadly, the problem isn't confined to Alzheimer's.

Whenever only a few people decide what is "right," the same pattern of stifled innovation will generally manifest itself not by design or from malice, but because it's hard for a small group to be 100% right on what works and what doesn't -- especially on matters as inscrutable as neuroimmune diseases.

TLDR: gatekeepers stifled exploration and innovation.

When a topic only has a limited number of experts, those experts become gatekeepers.

Those gatekeepers directly or indirectly control research funding.

Gatekeepers necessarily harbor biases, some right and some wrong, about how the field should progress.

For Alzheimer's, some gatekeepers were conflicted and potentially directed the field in the wrong direction. Only time will reveal AB42's true role.

It's easy to find fault in Alzheimer's.

It's harder to see the general solution to the gatekeeper problem, i.e., how to allocate resources in areas with limited experts.

VCs are soccer stars, but founders play basketball.

It’s easy to dunk on VCs, but the herd effect is rational after considering the typical VC’s background, the intense competition for good deals, and the job requirements — to prudently deploy capital.

Who wants to pitch their boss on investing $1-10M in a product no one uses, built by a team of anons?

This is not to defend the process, but merely explain it. It’s not so different from customer marketing. To win a VC, first understand the VC.

Once hired, VCs cannot easily get fired yet they exert immense strategic control.

Nonetheless, many founders interview summer interns harder than VCs.

Heuristic: after removing capital, would you hire the VC to be your boss?

Great VCs are worth the equity and will turbocharge startups. When you find one, don't haggle. Get a fair deal, and get right back to coding.

Bad VCs will destroy companies the same way soccer stars would destroy basketball teams if made the head coach.

The association between pathogens and cancer is under-appreciated, mostly due to limitations in detection methods.

For instance, it is not uncommon for cancer studies to design assays around non-oncogenic strains, or for assays to use primer sequences with binding sites mismatched to a large number of NCBI GenBank genomes.

Another example: studies relying on The Cancer Genome Atlas (TCGA), which is a rich database for cancer investigations. However, the TCGA made a deliberate tradeoff to standardize quantification of eukaryotic coding transcripts but at the cost of excluding non-poly(A) transcripts like EBER1/2 and other viral non-coding RNAs -- thus potentially understating viral presence.

Enjoy the rabbit hole. :)

A more accurate title: "Are Cornell Students Meritocratic and Efficiency-Seeking? Evidence from 271 MBA students and 67 Undergraduate Business Students."

This topic is important and the study interesting, but the methods exhibit the same generalizability bias as the famous Dunning-Kruger study.

The referenced MBA students -- and by extension, the elites -- only reflect 271 students across two years, all from the same university.

By analyzing biased samples, we risk misguided discourse on a sensitive subject.

@dang

Thanks. This is helpful. Looking forward to more of your thoughts.

Some nuance:

What happens when the methods are outdated/biased? We highlight a potential case in breast cancer in one of our papers.

Worse, who decides?

To reiterate, this isn’t to discourage the idea. The idea is good and should be considered, but doesn’t escape (yet) the core issue of when something becomes a “fact.”

Valid critique, but one addressing a problem above the ML layer at the human layer. :)

That said, your comment has an implication: in which fields can we trust data if incentives are poor?

For instance, many Alzheimer's papers were undermined after journalists unmasked foundational research as academic fraud. Which conclusions are reliable and which are questionable? Who should decide? Can we design model architectures and training to grapple with this messy reality?

These are hard questions.

ML/AI should help shield future generations of scientists from poor incentives by maximizing experimental transparency and reproducibility.

Apt quote from Supreme Court Justice Louis Brandeis: "Sunlight is the best disinfectant."

100% agreed. I also advise you not to read many cancer papers, particularly ones investigating viruses and cancer. You would be horrified.

(To clarify: this is not the fault of scientists. This is a byproduct of a severely broken system with the wrong incentives, which encourages publication of papers and not discovery of truth. Hug cancer researchers. They have accomplished an incredible amount while being handcuffed and tasked with decoding the most complex operating system ever designed.)

To elaborate, errors go beyond data and reach into model design. Two simple examples:

1. Nucleotides are a form of tokenization and encode bias. They're not as raw as people assume. For example, classic FASTA treats modified and canonical C as identical. Differences may alter gene expression -- akin to "polish" vs. "Polish".

2. Sickle-cell anemia and other diseases are linked to nucleotide differences. These single nucleotide polymorphisms (SNPs) mean hard attention for DNA matters and single-base resolution is non-negotiable for certain healthcare applications. Latent models have thrived in text-to-image and language, but researchers cannot blindly carry these assumptions into healthcare.

There are so many open questions in biomedical AI. In our experience, confronting them has prompted (pun intended) better inductive biases when designing other types of models.

We need way more people thinking about biomedical AI.

This is long overdue for biomedicine.

Even Google DeepMind's relabeled MedQA dataset, created for MedGemini in 2024, has flaws.

Many healthcare datasets/benchmarks contain dirty data because accuracy incentives are absent and few annotators are qualified.

We had to pay Stanford MDs to annotate 900 new questions to evaluate frontier models and will release these as open source on Hugging Face for anyone to use. They cover VQA and specialties like neurology, pediatrics, and psychiatry.

If labs want early access, please reach out. (Info in profile.) We are finalizing the dataset format.

Unlike general LLMs, where noise is tolerable and sometimes even desirable, training on incorrect/outdated information may cause clinical errors, misfolded proteins, or drugs with off-target effects.

Complicating matters, shifting medical facts may invalidate training data and model knowledge. What was true last year may be false today. For instance, in April 2024 the U.S. Preventive Services Task Force reversed its longstanding advice and now urges biennial mammograms starting at age 40 -- down from the previous benchmark of 50 -- for average-risk women, citing rising breast-cancer incidence in younger patients.

The author is a respected voice in tech and a good proxy of investor mindset, but the LLM claims are wrong.

They are not only unsupported by recent research trends and general patterns in ML and computing, but also by emerging developments in China, which the post even mentions.

Nonetheless, the post is thoughtful and helpful for calibrating investor sentiment.

More like alarming anecdote. :) Google did a wonderful job relabeling MedQA, a core benchmark, but even they missed some (e.g., question 448 in the test set remains wrong according to Stanford doctors).

For ML, start with MedGemma. It's a great family. 4B is tiny and easy to experiment with. Pick an area and try finetuning.

Note the new image encoder, MedSigLIP, which leverages another cool Google model, SigLIP. It's unclear if MedSigLIP is the right approach (open question!), but it's innovative and worth studying for newcomers. Follow Lucas Beyer, SigLIP's senior author and now at Meta. He'll drop tons of computer vision knowledge (and entertaining takes).

For bio, read 10 papers in a domain of passion (e.g., lung cancer). If you (or AI) can't find one biased/outdated assumption or method, I'll gift a $20 Starbucks gift card. (Ping on Twitter.) This matters because data is downstream of study design, and of course models are downstream of data.

Starbucks offer open to up to three people.

Thanks, but no one truly understands biomedicine, let alone biomedical ML.

Feynman's quote -- "A scientist is never certain" -- is apt for biomedical ML.

Context: imagine the human body as the most devilish operating system ever: 10b+ lines of code (more than merely genomics), tight coupling everywhere, zero comments. Oh, and one faulty line may cause death.

Are you more interested in data, ML, or biology (e.g., predicting cancerous mutations or drug toxicology)?

Biomedical data underlies everything and may be the easiest starting point because it's so bad/limited.

We had to pay Stanford doctors to annotate QA questions because existing datasets were so unreliable. (MCQ dataset partially released, full release coming).

For ML, MedGemma from Google DeepMind is open and at the frontier.

Biology mostly requires publishing, but still there are ways to help.

After sharing preferences, I can offer a more targeted path.

Agreed. There is deep potential for ML in healthcare. We need more contributors advancing research in this space. One opportunity as people look around: many priors merit reconsideration.

For instance, genomic data that may seem identical may not actually be identical. In classic biological representations (FASTA), canonical cytosine and methylated cytosine are both collapsed into the letter "C" even though differences may spur differential gene expression.

What's the optimal tokenization algorithm and architecture for genomic models? How about protein binding prediction? Unclear!

There are so many open questions in biomedical ML.

The openness-impact ratio is arguably as high in biomedicine as anywhere else: if you help answer some of these questions, you could save lives.

Hopefully, awesome frameworks like this lower barriers and attract more people.

To provide more color on cancers caused by viruses, the World Health Organization (WHO) estimates that 9.9% of all cancers are attributable to viruses [1].

Cancers with established viral etiology or strong association with viruses include:

- Cervical cancer - Burkitt lymphoma - Hodgkin lymphoma - Gastric carcinoma - Kaposi’s sarcoma - Nasopharyngeal carcinoma (NPC) - NK/T-cell lymphomas - Head and neck squamous cell carcinoma (HNSCC) - Hepatocellular carcinoma (HCC)

[1] https://pmc.ncbi.nlm.nih.gov/articles/PMC8831861

It's unclear why this drew downvotes, but to reiterate, the comment merely highlights historical facts about the CUDA moat and deliberately refrains from assertions about NVDA's long-term prospects or that the CUDA moat is unbreachable.

With mature models and minimal CUDA dependencies, migration can be justified, but this does not describe most of the LLM inference market today nor in the past.

Alternatives exist, especially for mature and simple models. The point isn't that Nvidia has 100% market share, but rather that they command the most lucrative segment and none of these big spenders have found a way to quit their Nvidia addiction, despite concerted efforts to do so.

For instance, we experimented with AWS Inferentia briefly, but the value prop wasn't sufficient even for ~2022 computer vision models.

The calculus is even worse for SOTA LLMs.

The more you need to eke out performance gains and ship quickly, the more you depend on CUDA and the deeper the moat becomes.

Google was omitted because they own the hardware and the models, but in retrospect, they represent a proof point nearly as compelling as OpenAI. Thanks for the comment.

Google has leading models operating on leading hardware, backed by sophisticated tech talent who could facilitate migrations, yet Google still cannot leap over the CUDA moat and capture meaningful inference market share.

Yes, training plays a crucial role. This is where companies get shoehorned into the CUDA ecosystem, but if CUDA were not so intertwined with performance and reliability, customers could theoretically switch after training.

Nvidia (NVDA) generates revenue with hardware, but digs moats with software.

The CUDA moat is widely unappreciated and misunderstood. Dethroning Nvidia demands more than SOTA hardware.

OpenAI, Meta, Google, AWS, AMD, and others have long failed to eliminate the Nvidia tax.

Without diving into the gory details, the simple proof is that billions were spent on inference last year by some of the most sophisticated technology companies in the world.

They had the talent and the incentive to migrate, but didn't.

In particular, OpenAI spent $4 billion, 33% more than on training, yet still ran on NVDA. Google owns leading chips and leading models, and could offer the tech talent to facilitate migrations, yet still cannot cross the CUDA moat and convince many inference customers to switch.

People are desperate to quit their NVDA-tine addiction, but they can't for now.

[Edited to include Google, even though Google owns the chips and the models; h/t @onlyrealcuzzo]

To address the downvotes, this comment isn't guaranteeing OAI's success. It merely notes the remarkably elevated probability of OAI escaping Nadella's grip, which was nearly unfathomable 12 months ago.

Even after breaking free, OAI must still contend with intense competition at multiple layers, including UI, application, infrastructure, and research. Moreover, it may need to battle skilled and powerful incumbents in the enterprise space to sustain revenue growth.

While the outcome remains highly uncertain, the progress since the board fiasco last year is incredible.