HN user

adidoit

239 karma

Founder @Socratify. Building a professional upskilling coach for the age of AI.

X/Twitter: @adidoit LinkedIn: https://www.linkedin.com/in/adipradhan/

Reach out hello@socratify.com

Posts4
Comments68
View on HN

With studies like these it's important to keep in mind selection effects.

Most of the high volume enterprise use cases use their cloud providers (e.g., azure)

What we have here is mostly from smaller players. Good data but obviously a subset of the inference universe.

This is fantastic. I'm reminded of the Samo Burja thesis that civilization is actually a lot older when we think and ancient civilizations including the Bronze Age were much more advanced than we think.

With better imaging, tooling, and archaeological funding, I'm sure we'll find much more evidence like this

So many countries bronze and ancient ages are underexplored

Yeah I think all of the concerns about ARPU and what the ROI from AI will be are not justified given the opportunity if executed well. LLMs contain high intent significant memory. Their usage is exploding.

Getting $200 subscriptions from a small number of whales, $20 subscriptions from the average white-collar worker, and then supporting everything us through advertising seems like a solid revenue strategy

Fascinating that the state-of-the-art in building agentic harnesses for long running agent workflows is to ... "use strong-worded instructions"

Anthropomorphism of LLMs is obviously flawed but remains the best way to actually build good Agents.

I do think this is one thing that will hold enterprise adoption back: can you really trust systems like these in production where the best control you can offer is that you're pleading with it to not do something?

Of course good engineering will build deterministic verification and scaffolds into prevent issues but it is a fundamental limitation of LLMs

This sounds like one of the "Ironies of Automation" as Lisain Bainbridge pointed out several years ago.

The more prevalent automation is, the worse humans do when that automation is taken away. This will be true for learning now .

Ultimately the education system is stuck in a bind. Companies want AI-native workers, students want to work with AI, parents want their kids to be employable. Even if the system wants to ensure that students are taught how to learn and not just a specific curriculum, their stakeholders have to be on board.

I think we're shifting to a world where not only will elite status markers like working at places like McKinsey and Google be more valuable but also interview processes will be significantly lengthened because companies will be doing assessments themselves and not trusting credentials from an education system that's suffering from great inflation and automation

The biggest challenge with LinkedIn is the primacy of your linkage to your employer enforces a self-censorship that turns into corporate speak.

The overton window on LinkedIn is actually quite small and because everyone there is really an employee rather than an employer, you get essentially slop that has been easily trained on and therefore is easily generatable by AI. It's just all low perplexity takes.

There's mostly no room for nuance because of the performative takes. Unlike a forum like Hacker News where your identity is almost totally abstracted away, every LinkedIn post is a move in the status game of career visibility.

LLMs have high utility for coding. The PMF is so strong that even with just coding i feel they will see an ARPU expansion beyond conservative assumptions.

There are adjacancies in white collar work like financial analysis that they will go after. All these will capture high ARPU usage.

Consumer is not their only path to revenue but it is probably the easiest to model. The enterprise play to automate and accelerate some white collar workers is a clear target not reflected here.

Claude Opus 4.5 8 months ago

Tested this building some PRs and issues that codex-5.1-max and gemini-3-pro were strugglig with

It planned way better in a much more granular way and then execute it better. I can't tell if the model is actually better or if it's just planning with more discipline

This is exactly why I'm focusing on job readiness and remediation rather than the education system. I think working all this out is simply too complex for a system with a lot of vested interest and that doesn't really understand how AI is evolving. There's an arms race between students, teachers, and institutions that hire the students.

It's simply too complex to fix. I think we'll see increased investment by corporates who do keep hiring on remediating the gaps in their workforce.

Most elite institutions will probably increase their efforts spent on interviewing including work trials. I think we're already seeing this with many of the elite institutions talking about judgment, emotional intelligence critical thinking as more important skills.

My worry is that hiring turns into a test of likeability rather than meritocracy (everyone is a personality hire when cognition is done by the machines)

Source: I'm trying to build a startup (Socratify) a bridge for upskilling from a flawed education system to the workforce for early stage professionals

It's very interesting of course at a natural move for aggregation. I'd love to see Stratechery / Ben Thompson's take on this.

From my perspective the challenges for vendors and SaaS providers are [1] discovery [2] monetization [3] disintermediation

I think it's less of a concern if you're Shopify or those large companies that have existing brand moats.

But if you're a startup, I don't think MCP as a channel is a clear-cut decision. Maybe you can get distribution but monetization is not defined.

Also I'm sure the model providers will capture usage data and could easily disintermediate you , especially if your startup is just a narrow set of prompts and a UX over a specific workflow.

The Reforge guys have been talking about a channel shift and this being it but until incentives are clear I'm not sure this is it yet. Maybe an evolution of this.

I'm building an AI coach for job seekers / early stage professionals (Socratify) and while I'd love more distribution from MCP UI integration I think at this point risk is higher than reward...

100% of AI startups are just multiplying matrices 100% of tech startups are just database engineering

It's still early in the paradigm and most startups will fail but those that succeed will embed themselves in workflows.

The older i get the more I see watterson's comics as a form of philsophical commentary on the state and evolution of the world

it's art and meaning and entertainment all in one

if it's been a while you should read them again with fresh eyes.

Google Antigravity 8 months ago

It's insane to me that I can't pay $20/$200 bucks after running out of limits in ~5 messages.

Why would you not at least link it to the pro and ultra accounts

at least you could upsell the pro subs to ultra. Millions of claude code and codex users who are into agentic coding is your servicable market paying attention today.

Now I'll delete antigravity and go back to codex / claude code / cursor ...

gemini cli. It's not as impressive as claude code or even codex.

Claude code seems to be more compatible with the model (or the reverse) whereas gemini-cli still feels a bit awkward (as of 2.5 Pro). I'm hoping its better with 3.0!

Yes I think specs as the context entry point is a great framing.

The word "spec" is a bit overloaded and I think we're all using it to define many things. There's a high-level spec and there are detailed component-level specs all of which kind of co-exist.

I think we're conflating spec-driven development with waterfall because most people's understanding of a spec is an extremely long document often with more bureaucratic components than actually useful information. It's the sort of thing that is done for legal reasons rather than engineering reasons.

However short and targeted specifications at the right level of detail and fidelity, can be extremely useful during coding with agents

One reason I haven't used Haiku in production at Socratify it's the lack of structured output so I hope they'll add it to Haiku 4.5 soon.

It's a bit weird it took Anthropic so long considering it's been ages since OpenAI and Google did it I know you could do it through tool calling but that always just seemed like a bit of a hack to me

I think OpenAI and all the other chat LLMs are going to face a constant battle to match personality with general zeitgeist and as the user base expands the signal they get is increasingly distorted to a blah median personality.

It's a form of enshittification perhaps. I personally prefer some of the GPT-5 responses compared to GPT-5.1. But I can see how many people prefer the "warmth" and cloying nature of a few of the responses.

In some sense personality is actually a UX differentiator. This is one way to differentiate if you're a start-up. Though of course OpenAI and the rest will offer several dials to tune the personality.

https://socratify.ai

Career Skills AI Coach. Sharpen how you think and speak by debating AI

We are clearly on the verge of the largest white-collar skills dislocation ever. Our goal at Socratify is to make skill building and reskilling for interviewing, onboarding, promotions, and career change as effective as possible with an AI coach and sparring partner.

Claude Skills 9 months ago

All of it is ultimately managing the context for a model. Just different methods

https://x.com/sayashk/status/1978565190057869344

AI agents have been developed for complex real-world tasks from coding to customer service. But AI agent evaluations suffer from many challenges that undermine our understanding of how well agents really work. We introduce the Holistic Agent Leaderboard (HAL) to address these challenges. We make three main contributions. First, we provide a standardized evaluation harness that orchestrates parallel evaluations across hundreds of VMs, reducing evaluation time from weeks to hours while eliminating common implementation bugs. Second, we conduct three-dimensional analysis spanning models, scaffolds, and benchmarks. We validate the harness by conducting 21,730 agent rollouts across 9 models and 9 benchmarks in coding, web navigation, science, and customer service with a total cost of about $40,000. Our analysis reveals surprising insights, such as higher reasoning effort reducing accuracy in the majority of runs. Third, we use LLM-aided log inspection to uncover previously unreported behaviors, such as searching for the benchmark on HuggingFace instead of solving a task, or misusing credit cards in flight booking tasks. We share all agent logs, comprising 2.5B tokens of language model calls, to incentivize further research into agent behavior. By standardizing how the field evaluates agents and addressing common pitfalls in agent evaluation, we hope to shift the focus from agents that ace benchmarks to agents that work reliably in the real world.

Is that true? I was under the impression that one reason NVIDIA is so entrenched is that their GPUs are highly flexible. Versus the ASIC offerings of other specialized players are more focused on transformers remaining the dominant architecture.

i made a half hearted attempt while building my startup - http://www.socratify.com

The use case was to build a knowledge graph to drive recommendations for the next best thing the user should learn.

After a few weeks of getting frustrated I went back to good old Postgres and writing a few tools for agentic retrieval.

It seems the agents are smart enough to traverse a database in a graph like manner if you provide them with the right tooling and context