HN user

sethkim

417 karma

Founder of Sutro (https://sutro.sh/).

seth@sutro.sh

Posts23
Comments70
View on HN
sethkim.me 3mo ago

Useful Black Boxes

sethkim
3pts0
sutro.sh 1y ago

The End of Moore's Law for AI? Gemini Flash Offers a Warning

sethkim
113pts75
www.skysight.inc 1y ago

Classifying aviation-related posts on Hacker News with SLMs

sethkim
10pts2
www.skysight.inc 1y ago

Generating 1M Synthetic Humans

sethkim
5pts0
www.skysight.inc 1y ago

Classifying aviation-related posts on Hacker News with SLMs

sethkim
4pts0
sethkim.me 1y ago

Non-Scalar Leverage

sethkim
1pts0
www.skysight.inc 1y ago

Model Security with Large-Scale Inference

sethkim
3pts0
sethkim.me 2y ago

Magical Chasms

sethkim
1pts0
sethkim.me 2y ago

Order from Chaos: The Subtle Superpowers of Transformer Models

sethkim
2pts0
tailgate.dev 2y ago

Show HN: tailgate – Build generative-AI features without a backend

sethkim
1pts0
regrail.io 3y ago

Show HN: Regrail – Open-Source, Visual Data Transformation and Querying Tool

sethkim
6pts2
sethkim.me 3y ago

The Code Less Stupid

sethkim
1pts0
sethkim.me 3y ago

The Best Build Tools

sethkim
1pts0
news.ycombinator.com 3y ago

Ask HN: Does anyone want to take ownership of roundtable.audio?

sethkim
28pts11
sethkim.me 4y ago

The (Non-)Utility of College

sethkim
1pts0
www.forbes.com 4y ago

Contrary Raises $75M Fund III to Bet on the Founders of Tomorrow

sethkim
1pts0
sethkim.me 4y ago

Doors That Open Doors – What Is (and Isn't) a Technology Company?

sethkim
2pts0
sethkim.me 4y ago

The metaverse will be paved with new communication, not payments layers

sethkim
11pts14
sethkim.me 4y ago

Automation != Leverage

sethkim
58pts28
mantis.chat 4y ago

SaaS on a Hunch

sethkim
2pts0
mantis.chat 4y ago

Show HN: Mantis is an embeddable business phone in the browser

sethkim
14pts6
hackernews.roundtable.audio 5y ago

Show HN: hackernews.roundtable.audio turns HN posts into live audio discussions

sethkim
109pts65
news.ycombinator.com 5y ago

Show HN: Hacker News Discourse turns HN stories into Clubhouse-style audio rooms

sethkim
36pts11

We build a product that's somewhat similar in spirit to DSPy, but people come to us for different reasons than the OP listed here.

1) It's slow: you first have to get acquainted with DSPY and then get hand-labeled data for prompt optimization. This can be a slow process so it's important to just label cases that are ambiguous, not obvious.

2) They know that manual prompt engineering is brittle, and want a prompt that's optimized and robust against a model they're invoking, which DSPy offers. However, it's really the optimizer (ex. GEPA) doing the heavy-lifting.

3) They don't actually want a model or prompt at all. They want a task completed, reliably, and they want that task to not regress in performance. Ideally, the task keeps improving in production.

Curious if folks in this thread feel more of these pains than the ones in the article.

Under-discussed superpower of LLMs is open-set labeling, which I sort of consider to be inverse classification. Instead of using a static set of pre-determined labels, you're using the LLM to find the semantic clusters within a corpus of unstructured data. It feels like "data mining" in the truest sense.

My two cents here is the classic answer - it depends. If you need general "reasoning" capabilities, I see this being a strong possibility. If you need specific, factual information baked into the weights themselves, you'll need something large enough to store that data.

I think the best of both worlds is a sufficiently capable reasoning model with access to external tools and data that can perform CPU-based lookups for information that it doesn't possess.

No doubt prices will continue to drop! We just don't think it will be anything like the orders-of-magnitude YoY improvements we're used to seeing. Consequently, developers shouldn't expect the cost of building and scaling AI applications to be anything close to "free" in the near future as many suspect.

Both great points, but more or less speak to the same root cause - customer usage patterns are becoming more of a driver for pricing than underlying technology improvements. If so, we likely have hit a "soft" floor for now on pricing. Do you not see it this way?

I run a batch inference/LLM data processing service and we do a lot of work around cost and performance profiling of (open-weight) models.

One odd disconnect that still exists in LLM pricing is the fact that providers charge linearly with respect to token consumption, but costs are actually quadratic with an increase in sequence length.

At this point, since a lot of models have converged around the same model architecture, inference algorithms, and hardware - the chosen costs are likely due to a historical, statistical analysis of the shape of customer requests. In other words, I'm not surprised to see costs increase as providers gather more data about real-world user consumption patterns.

Sutro.sh (fka Skysight) | Infrastructure/LLMs & Research Engineering | SF Bay Area | Full-time

We are building batch inference infrastructure and a great/user developer experience around it. We believe LLMs have not yet been meaningfully unlocked as data processing tools - we're changing that.

Our work involves interesting distributed systems and LLM research problems, newly-imagined user experiences, and a meaningful focus on mission and values.

Open Roles:

Infrastructure/LLM Engineer — https://jobs.skysight.inc/Member-of-Technical-Staff-Infrastr...

Research Engineer - https://jobs.skysight.inc/Member-of-Technical-Staff-Research...

If you're interested in applying, please send an email to jobs@sutro.sh with a resume/LinkedIn Profile. For extra priority, please include [HN] in the subject line.

Skysight | Infrastructure/LLMs & Research Engineering | SF Bay Area | Full-time

We are building large-scale batch inference infrastructure and a great/user developer experience around it. We believe LLMs have not yet been meaningfully unlocked as data processing tools - we're changing that.

Our work involves interesting distributed systems and LLM research problems, newly-imagined user experiences, and a meaningful focus on mission and values.

Open Roles:

Infrastructure/LLM Engineer — https://jobs.skysight.inc/Member-of-Technical-Staff-Infrastr...

Research Engineer - https://jobs.skysight.inc/Member-of-Technical-Staff-Research...

If you're interested in applying, please send an email to jobs@skysight.inc with a resume/LinkedIn Profile. For extra priority, please include [HN] in the subject line.

Skysight | Infrastructure/LLMs & Product Engineering | SF Bay Area | Full-time

We are building large-scale, data-intensive inference tooling and a great/user developer experience around it. We believe LLMs have not yet been meaningfully unlocked as data processing tools - we're changing that.

Our work involves fascinating distributed systems and ML research problems, newly-imagined and well-crafted user experiences, and a meaningful focus on mission and values.

Open Roles:

Infrastructure/LLM Engineer — https://jobs.skysight.inc/Member-of-Technical-Staff-Infrastr...

Product Engineer - https://jobs.skysight.inc/Member-of-Technical-Staff-Product-...

If you're interested in applying, please send an email to jobs@skysight.inc with a resume/LinkedIn Profile. For extra priority, please include [HN] in the subject line.

What's extremely confusing to me (as a private pilot) is that traffic is almost always routed directly over an airport (midfield), to safely avoid departing and landing traffic. The sense that I get is that it became routine for traffic to be routed directly through the glidepath in a staggered manner, likely because it's military.

Such unsafe habits (like driving without a seatbelt on) statistically will eventually result in a tragic outcome.

This is really cool, and hints at a near-future possibility of building a search engine on top of just about anything. It's clear we've moved past the ability to just search for website url's and webpage content. Anything that can be indexed - regardless of type of data or dimension (space, time, etc.) will be searchable.

I figured this comment would get me in trouble :)

I recommend doing some instrument lessons if you haven't already. When I got my instrument rating I questioned whether the private requirements are actually enough. The skills that the instrument rating teaches in terms of preparation, workload management, and emergency weather scenarios make me question whether I was really ready to fly before I had it.

With the private or sport license you'll be fine in the majority of cases. I think my comment comes more from the edge and corner cases that more skills and experience help with, not your ability to work in the system at all.

In single pilot IFR, an autopilot is often your best friend. It's exactly like you say - when you're busy with everything else you want the plane to fly itself. Isn't that problem already somewhat solved in a sense? Or are you referring to Garmin Autoland (or similar) in emergencies?

By the way - I can totally see how a great GA fly-by-wire system is an improvement to maintain positive control of an aircraft at all times. I'd personally love to give it a try and see how it reduces pilot effort while flying.

What if this tech made the individual flyer safer?

I'd hope that's the case! That's why I put it in the "good" category".

how much could be in the air at one time?

Hard to say, but there's a ton of congestion around busy airspace as is. I'd think an order-of-magnitude increase in GA traffic would require a major rework of the whole airspace system.

Instrument-rated pilot (and engineer) here.

First - congrats on the launch! I think you're working on an interesting set of components that will prove useful to GA aircraft technology. Bringing fly-by-wire, and lowering the cost of maintenance/manufacturing are both great efforts.

That being said, my personal view is that stick-and-rudder control is one of the less critical components to improving GA safety. Everything else - flight planning, comms, automation, navigation, weather, inspections, procedures, regs, and most importantly - working in the federal airspace system - are the "hard" parts of flying and where problems tend to occur. It's common belief that single-pilot IFR is the most challenging type of flying, because of how much you have to do all at once.

It may sound snobby - but I'm not super excited about the idea of lowering the barrier to entry for GA on a foundational skill basis. Like the light-sport rating, it encourages more people to be in the (already congested) airspace system who haven't really gained all the other skills necessary or experience to be there.

To be clear - I think improving technology and lowering costs = good. Lowering early-skill requirements for pilots and pushing more people without all the other skills into federal airspace = very bad. In general, I'd frame this effort more as an effort to raise the bar for system technology, not lower the bar to become a pilot in the first place.

It became clear to me in the GPT-4o launch that OAI was interested in the "GitHub copilot for everything" route. With low-latency voice mode and the ability to take action on a user's behalf, this will basically feel like pair-programming for everything, or just having an employee who can do everything for you until they need your help or clarification.

To power that experience, an app will need to feel "multiplayer" like someone else is working with you. They'll probably bundle this in the API and have "agent mode" that developers can embed in any app or website, or just let consumers give OAI access to control their desktop. It'll also likely work async, so you can assign tasks, walk away or go to sleep for a few hours, and see the results.

This is speculation. But it feels like the interface that we'll look back on in ten years and say "that seemed obvious in hindsight."

What's cool is that Erik actually acted on these complaints. Modal is, by far, my favorite developer tool ever and makes me hopeful not just for the future of software engineering but the entire tech industry.

If you're a naysayer in the comments, I would encourage you to go give it an honest try, and consider again why you think infra has to be done in harder ways.

HTML First 3 years ago

I built something recently with the same ideas in mind: https://github.com/sethkimmel3/tailgate. It allows people to build generative-AI applications without any of the complication of setting up a backend, and in the simplest cases only requires adding HTML attributes.

I originally learned to code with HTML, CSS, and JS, and I think it's still the easiest way to experience the magic of shipping a working application to others. We should keep encouraging more patterns and tooling that lower the barriers of entry to those just starting out.

My two cents is that we will see a bubble pop, but it will rebound massively much like SaaS following the dotcom crash.

My primary concern is the ratio of AI-infrastructure products being built relative to the average utility delta of current end-user products. There are some shining counterexamples to this like Github copilot, but for the most part the utility gain from most text-based generative-AI seems incremental.

Current language models are bad to okay at most tasks - and great at a select few. The capabilities in unstructured ETL and production of labeled datasets fall in the latter category.

I think the next generation of really powerful capabilities to be unlocked with generative-AI will come with 1) an increase in compute availability (training even larger models) 2) advances in multimodal models, and 3) better interfaces to interact with them. Spatial computing will be a really big deal here, and people seem to have forgotten that Apple may have unlocked the next version of interactive computing earlier this year.

I have a really hard time understanding how this, and similar products like Scale Donovan make sense in a defense capacity. Even more so than previous AI waves, LLM's are highly non-deterministic and are way more suited to low-risk use cases (for now). I understand the appeal of giving more data-driven capabilities to non-analytical users, but when it risks returning hallucinated results that could lead to actual human casualties - that quickly crosses a line.

I have a ton of respect for both Palantir and Scale as companies - but this just looks like a cash grab race to see who can dupe the Pentagon the fastest into believing they need some technology that's years away from anything they should be using in practice.

It seems like Hacker News might be broken, since all of the top posts have like one or two points. Anyway, seems like a good opportunity to help someone out!

Everyone has their own value function in life, and it will change over time. My advice is generally to do what feels important. I wrote about this recently: https://sethkim.me/l/thesolutionspace/?postnum=7.

If the internship opportunity seems incredible, it probably feels important to do, right? And I'd imagine if the pay is really good, you'll be able to travel home a handful of times over the summer and spend time with family/friends. But if spending the whole summer with them feels more important than the internship, it probably is. It's your question to ask yourself, and answer honestly.

On the question of growing up too fast: you will most likely have to get a start on the career you think you want at some point, so the internship sounds like a smart move. I say "you think", because that's likely to change as well (sorry to say). Internships are a great way to dip your toe in the water of a number of career paths to start to figure that out.

Life is a great balancing act of doing what you have to do, should do, and want to do. Balance wisely, but don't forget to have fun along the way!