HN user

alansaber

435 karma
Posts23
Comments391
View on HN
alanyahya.com 5d ago

Domain Specific Harnesses

alansaber
2pts0
lexifina.com 7d ago

Legal AI, not a coding agent with scaffolding

alansaber
10pts0
lexifina.com 10d ago

The logic behind Kirkland x Palantir

alansaber
2pts0
lexifina.com 26d ago

Compaction in CC, Codex, and OpenCode

alansaber
2pts0
lexifina.com 27d ago

Multi agent systems for complex tasks

alansaber
2pts0
lexifina.com 27d ago

Building Agent Telemetry for LLMs

alansaber
3pts0
lexifina.com 1mo ago

Multi doc agent workflows in Word

alansaber
1pts0
lexifina.com 1mo ago

The Legal Agent Stack

alansaber
2pts0
alanyahya.com 1mo ago

"Cursor for X": key standards for vertical products offering agent workflows

alansaber
2pts0
lexifina.com 1mo ago

Orchestration problems hurt legal AI

alansaber
1pts0
news.ycombinator.com 2mo ago

Show HN: 3x Enemies Aggro Mod – Elden Ring Multiplayer Scaled Up to the Max

alansaber
1pts0
alanyahya.com 2mo ago

The bull case for graph DBs in law

alansaber
9pts2
lexifina.com 4mo ago

The Battle for the Billable Hour

alansaber
1pts0
alanyahya.com 4mo ago

Automated Materials Design

alansaber
1pts0
lexifina.com 4mo ago

Version control and full session capture are table stakes

alansaber
1pts0
lexifina.com 4mo ago

Will Claude Code Consume Legaltech?

alansaber
2pts0
arxiv.org 5mo ago

Task Specific Knowledge Graphs

alansaber
1pts0
lexifina.com 5mo ago

Visualising legal memory through knowledge graph diffs

alansaber
2pts0
lexifina.com 5mo ago

Ontologies are all you need

alansaber
3pts0
lexifina.com 7mo ago

A roadmap to build the SoTA for RAG

alansaber
1pts0
lexifina.com 7mo ago

Tech for Small vs. Big Firms

alansaber
15pts11
lexifina.com 8mo ago

Version Control for Lawyers

alansaber
2pts0
lexifina.com 9mo ago

Reinforcement Learning for Law

alansaber
2pts0

I never thought i'd see the day they released a model, rather than a blog post. The Figure 3 demo being a screencap of chrome in localhost made me feel better about myself. Jokes aside, best western open weights model- very cool.

Interesting. New models are estimated at ~5T params, so 45,000x increase over BERT base (110m). But vocab size of 200k, so only an increase of 7x over BERT base (30k).

A lot of people want a use case. One I think might be cool is some kind of spatial/represented comparison: let's see how two different models interact with the codebase (for the same problem), what they touched, and what they did. Or the same model, but averaged across 100 runs, so we can see how much variance there really is per task. Something along those lines sounds interesting to me.

clearly benchmark and optimise for a specific model over millions of datapoints > new model comes out > get to do it all over again. At this point just become Cursor and get paid for it.