HN user

zurfer

811 karma

Building https://getdot.ai to chat with your data stack.

Posts22
Comments293
View on HN
status.claude.com 4mo ago

Increased Errors on Opus 4.6

zurfer
3pts0
status.claude.com 4mo ago

Elevated errors on Claude Opus 4.6 and Sonnet 4.6

zurfer
3pts1
status.claude.com 4mo ago

Elevated errors on login with Claude Code

zurfer
60pts58
iss-sim.spacex.com 1y ago

SpaceX – ISS Docking Simulator

zurfer
2pts0
openai.com 1y ago

OpenAI o1 API and new tools for developers

zurfer
9pts1
www.kaggle.com 1y ago

3rd Place Solution for the Arc Prize 2024 Competition: Omni-Arc Approach

zurfer
1pts1
news.ycombinator.com 1y ago

Ask HN: Are models changing for pinned releases?

zurfer
2pts1
01-ai.github.io 2y ago

01.ai Yi-Large LLM Launch

zurfer
1pts1
hn-with-topics.web.app 2y ago

Hacker News with headline-based topic extraction

zurfer
2pts2
github.com 2y ago

GPT-4-turbo-2024-04-09 "wins" simple evals benchmark

zurfer
2pts0
github.com 2y ago

A list of AI analytics tools

zurfer
4pts0
hub.zenoml.com 2y ago

Gemini Benchmark – MMLU (compared with GPT-4-turbo, Mixtral)

zurfer
1pts1
eu.getdot.ai 2y ago

Hacker News Activity Analysis with GPT-4 Agent

zurfer
139pts51
news.ycombinator.com 2y ago

Ask HN: Why is GPT-4 or Claude-2 so bad at tic-tac-toe?

zurfer
2pts7
github.com 2y ago

Microsoft drops support for Python SDK for Teams

zurfer
1pts1
github.com 2y ago

SQLite-Vss: A SQLite Extension for Vector Search

zurfer
4pts1
status.openai.com 3y ago

OpenAI Outage

zurfer
130pts154
github.com 3y ago

Open Source Multi-Modal GPT

zurfer
2pts0
www.metabase.com 4y ago

How to Document Data

zurfer
3pts0
www.snowboard.software 4y ago

Data for everyone and more bad ideas

zurfer
4pts1
news.ycombinator.com 5y ago

Ask HN: Is There a Stripe Atlas Alternative for EU or Germany?

zurfer
5pts7
www.youtube.com 5y ago

Stonebraker on databases and venture capital (2014)

zurfer
1pts1

more interesting link: https://arxiv.org/html/2607.09424v2 and > Long-context serving efficiency. Soofi S combines frontier-level capability with the highest measured aggregate long-context decode TPS, and unlike full-attention dense baselines maintains high throughput as context grows. Panel (1(a)) plots Capability Index versus measured aggregate decode TPS/GPU at 40K context and batch 32. The Capability Index averages five benchmark groups, i.e., Code, GSM8K, GPQA-Diamond, English aggregate, and German aggregate, after normalizing each group to the best plotted model. Aggregate decode TPS/GPU is measured with a TP=1, one-B200 vLLM latency-subtraction protocol. Panel (1(b)) shows measured aggregate decode TPS/GPU as a function of input context length under the same batch-32 protocol.

it's a small win in the small model class

Somehow the blog post seems naive. Yes GLM 5.2 is good and cheaper per token, but margins are a result of supply and demand. Now demand for quality and quantity of tokens is increasing at least quadratic or cubic (more users * more tasks * more tokens per task). On the other side you have real infrastructure constraints on the supply side. Openai and Anthropic have large commitments and contracts that enable them to get access at a scale of compute that is not obviously going to be available for open source model hosts. And you see it, glm 5.2 inference is less stable and higher variance than any of the bigs labs.

Why is SpaceX not hosting glm 5.2? because they make more money with renting out to Anthropic and Google.

So he throws billions at a few top AI researchers, but they produce nothing of value.

so he spends 1% of yearly revenue on AI talent to catch up? we can't judge if they have produced nothing of value, no? They don't owe the world to open source their work?

Meta has plenty of failings, but taking risks and investing optimistically is not on my list. I guess the sentiment here on HN is probably biased by the addictive nature of its products.

This is called mechanistic interpretability. There is lots of fascinating insights already since you can do basically everything down to the neuron or weight level thousands of times. The human brain is many orders of magnitude harder to make sense of.

We outsourced it for 2.5k (extra) and it was still painful, took almost 2 months and worst of all wasted so much time and focus.

The worst was sitting at the notary, and getting read out loud by her what we were about to sign (also paying for that).

If you think about starting a company, spend some time to think through what it would mean for you to be a Delaware C Corp or an Estoinian one. It will increase your chances of success as you can focus on what matters.

Like sqlite, duckdb is underappreciated as a production database. You can totally run it on servers or even "serverless" and do some heavy data transformations or with the right server size work with large scale datasets (up to a TB compressed seems fine).

Midjourney Medical 1 month ago

I heard the same argument from my doctor when I wanted a blood scan.

But what's the intention? If you do a scan and then try to find everything that is wrong about you, you're 100% right, there will be false positives and unnecessary panic/medication etc.

However if you just collect data for months and years and WHEN you get a symptom you have a lot more data then we should be able to give better diagnosis faster. If we do that for long enough as humanity and there is data sharing the accuracy of the whole thing will increase a lot.

I do think it's more subtle. AI can replace very few jobs end to end with the same quality, maybe none. But AI can be put to work on high ROI problems. Now when the new marginal job is not obviously as high ROI as putting another 100k of tokens to work, no human gets hired.

Next, comes natural attrition in a company where a certain percentage will leave every year. Will they get replaced with a human or their budget goes into tokens?

Only when these 2 angles are exhausted, a typical company will start thinking about layoffs.

Now, some companies are already stressed: customer buy AI products instead of theirs, AI makes it easier to build what they offer, customers believe they can vibe code things. These companies will layoff first, because of AI. Not because AI will do the persons job but because the money gets spend differently.

Techno optimist here who expects the following to make a big contribution to reducing human made future climate change: better batteries+solar/wind, nuclear fusion, self driving cars (we'll need to manufacture less cars for the same amount of miles humanity drives), AI helping with better resource allocation in general (hopefully).

The answer can't be, let's just consume 10x less. We have to engineer our way out of it.

It's priced at 1/10, but deepseek is probably not profitable, also it's slow.

Even more interesting is the question if we would have a deepseek model without the US frontier models.

And then what's the value of the advantage that the frontier models have. It's definitely 100x more valuable to find zero days 3months earlier. Probably not in every domain but in enough domains having the smartest model is valuable.

DeepSeek v4 3 months ago

lots of great stuff, but the plot in the paper is just chart crime. different shades of gray for references where sometimes you see 4 models and sometimes 3.

this is vibe coded :D. the drop down for Python not working.

I remember when we started working on our teams bot in 2023, Microsoft announced that they will stop supporting Python for the teams sdk, which felt super short sighted. Eventually they silently picked it up again and the old sdk never stopped working.

Kudos for turning it around. I believe it'll last this time as AI agents in communication tools are something a lot of people want. and unlike pre LLM chatbots and agents, they are so much more useful now.