Sort of! We came up with a backronym just in case: Scenario Oriented Reasoning and Outcome Simulations.
HN user
muggermuch
Blog: https://principiamundi.com/ Email: am@principiamundi.com
Building Lookback Labs: the intelligence layer for AI-native hedge funds.
Previously founded Didact AI, an ML-driven stock picking engine.
I'm often found quietly muttering about deep learning, financial markets, and computer science.
--- meet.hn/city/us-San-Francisco
Socials: - linkedin.com/in/anshumanmishra
A few days ago, we soft-launched Soros [0], our agentic intelligence for geopolitical simulation, on HN. [1]
In this post, we walk through everything: how we designed it, the engineering stack, how it works, what we perceive as flaws. Everything.
This is really a chance to look at how a production-grade agentic system works when designed for unconventional use cases.
Note that this is a dense, super-long article (clocking in at ~14,400 words). We suggest you take your time reading it :)
We look forward to your questions and comments!
- Anshuman & Karén, co-founders of Lookback Labs (creators of Soros) [2]
You can also reach us at: team@lookbacklabs.com
because this is marketed as being for macro investing, I would expect to see a level of rigor and quantitative analysis consistent with that.
Thanks for bringing this up - while we talk about Soros' forecasts and comparing them against those of an LLM, in the end Soros is not a forecasting tool, it's an analytical framework.
There is a gap between quant modeling and geopolitical analysis that we seek to fill. Specifically, quant models are great at capturing statistical regularities in financial time series but typically treat geopolitical shocks as exogenous noise. Meanwhile, geopolitical analyses in the policy and intelligence communities (with the exception of Bueno de Mesquita [BdM]'s work) provide deep contextual reasoning but rarely produce probabilistic scenario structures or asset-level transmission mappings that can directly inform capital allocation.
We will be shortly publishing a technical preprint laying out the Soros framework in full, but the TL;DR is: we model geopolitical events (or crises in the literature) as partially observed ("fog of war") stochastic games with multiple actors jostling for control over resources. We map out actors across various axes (think of these as actor embeddings), identify key decision points, and enumerate paths across them to estimate scenario probabilities. The scenarios in turn have associated transmission flows and market implications. We will evaluate those as mentioned in the sibling comment. Happy to discuss more.
First, thank you so much for signing up to try out Soros!
You are absolutely right, of course, to ask about accuracy. TL;DR: we don't have any formal calibration data yet.
The reason why is interesting, though, and it strikes at the heart of global macro investing in particular: things change, often, and sometimes dramatically. Basically, geopolitical "events" are really smeared across time (and sometimes space). Each event update can lead to a cascade of new scenarios branching off and older ones dying out, each with implications on capital flow. It's difficult to disentangle, which is why our preference has been to enable the system itself to monitor feeds, but also update its alerts as it deems fit, and re-run the analysis when it feels there's been enough of a change of state (pun not intended).
One markets-focused eval we have been building towards (and apparently you have been thinking of as well) is comparing against LLMs. Our plan is to run simultaneous comparisons against a variety of frontier models, armed with the same information that we provide Soros, but without the structural framework and simulation engine we've built though. Ideally we want to map out the Pareto frontier of model capability vs realized returns, and examine performance over horizons, asset classes, and so on, and have concrete numbers on where Soros pushes the curve outwards.
This is being built :), and we hope to get there in the coming few weeks!
Exactly! That's how it works - the static demo is just that, static.
Would love to onboard you for the full thing if you'd like! Just LMK (team@lookbacklabs.com) or add your info on the site
So someone monitoring Polymarket could have reached the same conclusion?
Maybe? If they are professionally trading prediction markets, I'm pretty sure that would be the case. Polymarket especially is a great source of insider traded information, as you pointed out.
We do near realtime tracking of most major markets, plus X accounts that Soros identifies as being important. The system also composes search queries per analysis, along with frequency of scanning, and that's run as requested. (We use a mix of Perplexity and other smaller search providers, along with Exa via OpenRouter's integration.)
Hope this helps! Thanks for your questions!
Hmmm, good question. I think one interesting incident for us was when we saw scenario probabilities being updated near last Friday EOD for the US-Iran conflict, biased towards further kinetic action by the US around Kharg island (?). This was basically captured from changes in odds for Polymarket events that the system was tracking. The news came in a few minutes later, post equity market closing.
* This brings us to a larger question - why did we build Soros?
First, let's address the elephant in the room: we were inspired by George Soros' theory of reflexivity and how human tendencies affect markets more prominently than expected. Yes, there's a corny backronym [0]. No, this is not a political statement or endorsement of his views.
Coming back to the main point, we (the founding team at Lookback Labs) have both spent a long time at the intersection of financial markets, technology, and machine learning. During that time, one key thing that kept bothering us [1] was simply this: when a geopolitical crisis breaks, an investor's actual problem is not really to find out "what is happening now" — it's more of "which scenario plays out, how likely is each one, and what do I buy, sell, or hedge under each? For how long?"
There are a ton of existing tools and services that seek to answer the first question reasonably well (newsletters such as StratFor, publications such as Foreign Affairs and Foreign Policy, Bloomberg terminals for breaking news, etc.).
None of these answer the other questions particularly deftly. Sure, one can engage with ChatGPT (or Claude if one prefers), and play through multiple scenarios. You will, of course, miss out on the grounded structural model that powers Soros' analysis, along with the simulations that serve up the relative probability estimates.
Also, one of the worst things purely LLM-based ad hoc frameworks do is assume that countries are monolithic decision-making units from a game-theoretic perspective. This is hardly the case - "Iran" doesn't make choices, Mojtaba and the IRGC faction does. "China" doesn't decide, the Politburo Committee does. And so on.
There are of course formal analytical frameworks that dig deeper, studying groups, factions, organizations that are jostling to gain control (Bruce Bueno de Mesquita's Expected Utility Model and selectorate theory [2] is the most academically serious and is a prime inspiration for our system design), but they are extraordinarily hard to operationalize in real time, and produce no market implications.
To sum up, the choices are stark: ask AI and hope for the best, or build out your own systematic framework to organize evidence, assumptions, and implications. We chose the latter path.
Zooming out, our mission at Lookback Labs (https://www.lookbacklabs.com/) is to build "the intelligence layer for AI-native investing"; accordingly, Soros is the first of several agentic systems that we are designing across the systematic and discretionary spaces, that are both usable and useful from the get go, and not merely demo eye candy.
* Some minor details:
(1) We are currently in private beta for Soros and are onboarding selectively.
(2) The static demo is not completely static; you can still chat with the analysis (up to 20 messages a day per IP).
(3) We are still working on pricing: something that captures the value Soros provides.
(4) We want this to work for individual investors as well, not just institutional desks, and would love to price accordingly.
We're curious to hear what the HN community thinks about our approach. AUA!
Feel free to reach out offline if you'd like! We are, sadly enough, on LinkedIn, but are also available via email (anshuman/karen@lookbacklabs.com)
PS: As is probably obvious to the diligent reader :), every token in this post has been lovingly handcrafted by the Lookback Labs team.
[0] Scenario-Oriented Reasoner for Opportunity Synthesis. Lol.
[1] Many things bothered us. Buy us drinks, get stories.
[2] We heartily recommend two of BdM's books: "Predicting Politics" and "The Dictator's Handbook"
Absolutely heart-warming to see some of the gentle tech titans of yesteryear still actively releasing products.
This scenario is oddly terrifying.
This is amazing - just the tool I needed; thank you so much!
Cool anecdote, thanks for sharing!
Good to see you here, Dave! Glad you decided to open source LLM Catcher.
I like this a lot!
But: I feel the more of these services come to being, the more likely it is that every website starts putting up gates to keep the bots away.
Sort of like a weird GenAI take on Cixin Liu's Dark Forest hypothesis (https://en.wikipedia.org/wiki/Dark_forest_hypothesis).
(Edited to add a reference.)
As a Harper's and Lapham's Quarterly subscriber, I have been a huge fan of his quirky editorial style.
Specifically, I'd like to call out his podcast ("The World In Time"). Its past episodes remain treasure troves of wisdom, with LL's resonant voice asking the kind of engaging questions that are a rarity these days. Highly recommended.
in Oshiro, i had met a guardian of time, a man who, year after year, preserved a slice of japan's essence, ensuring that even in the heart of its busiest city, the song of the bush warbler would never fade away.
There was a lump in my throat as I read your comment out loud to my wife. Thank you for sharing this beautiful vignette!
As per the Massive Text Embedding Benchmark (MTEB) Leaderboard maintained by Huggingface, OpenAI's embedding models are not the best.
https://huggingface.co/spaces/mteb/leaderboard
Of course, that's far from saying that they're the worst, or even headed that way. Just not the best (those would be a couple of fully opensource models, including those of the Instructor family, which we use at my workplace).
It's changed my life.
A few months ago, I wrote a post [0] talking about my experiences designing and building an ML-powered stock picking engine for my startup - the post went viral on HN, and it led to many fascinating conversations, valuable connections, opportunities to speak, and job offers (tech/ML, tradfi, and crypto). In fact, it quite directly led to my new job, as a team reached out with an opportunity that ticked off all the boxes I was looking for.
Finally, thanks to my blog, I have made many new friends who I hope to engage with productively[1] going forward, and I feel more firmly embedded in the intellectual milieu of the Bay Area than I ever did. As a consequence, I am much more relaxed now and feel in control of the overall direction of my life.
[0] https://principiamundi.com/posts/didact-anatomy/
[1] The meaning of this may change over time
I have been following your blog for a while now - it's been quite instructive on ways to speed up LLM-based tasks. All the best with your launch!
This looks extremely useful! Thanks a ton.
1. I would fiercely recommend reading “An Engine, Not A Camera” (Donald MacKenzie). Formally, it reads as a sociological study of the practice of finance - how financial markets shape society instead of simply capturing its needs and desires at large. More importantly, however, it made me think deeper on questions about fields of study that are (were?) distant and to question assumptions: a lot of ideas that we take for granted as absolute truths are simply consensus-derived and have no objective reality.
2. On a similar theme, “The Lady Tasting Tea” (David Salsburg) was a wonderful history of the development of statistics as a mathematical discipline. I found it absolutely fascinating to map out how individual personalities were buffeted and shaped by larger historical events and movements (e.g. eugenics in the late 19th century, WW2, the rise of the Soviet Union, etc.) into asking questions of data that propelled statistics forward. Disciplines we think of as being “hard” (math, CS, statistics, physics) have been shaped by social forces in a non-linear non-Hegelian fashion, in fits and starts, quite contrary to the way they are presented in a pedagogical setting.
FTX's books don't so much have red flags as that they are printed on red flags, with ink derived from red flags, a custom-made red cover made out of more red flags, and each page, when opened, has pop-up red flags along with a little electronic speaker that plays Red Flag by Antigoni.
Lol. This should be a framed quote.
This is very useful; thank you for the wonderful blog post!
This is a very insightful remark, thank you.
I focused on 10-Qs for the EDGAR filings module as you rightly pointed out - it seemed to be a good balance between implicit information and usefulness of the data. TBH I didn't actually investigate the other (many) patterns.
Having said that, I have really enjoyed Kai Wu's research from Sparkline Capital (https://www.sparklinecapital.com/), especially his extraction of the innovation factor from EDGAR filing texts. He's appeared in numerous podcasts, and they have all been super useful to listen to. Maybe someday when I re-investigate EDGAR filings and go further, I might target these signals you talk about here.
Thank you, and ha! An emphatic yes to all the points you raised! It's especially daunting when there are multiple vendors with incompatible point-in-time hygiene setups, a situation I faced at the beginning of setting up Didact.
Also, this was really my first time with equities - my professional trading career was derivatives-focused - both listed (CME) and OTC (FX forwards/swaps). I think I lost the first few months simply trying to reorient my style of thinking.
Water under the bridge I guess.
Thank you for the note! Just picked up Isichenko from the online bookstore we all love-hate.
I'd love to get in touch (as per your HN profile) - my email is am(at)principiamundi.com
:) Will do! Thanks for the encouragement!
Yes! I use fractional Kelly extensively in my (separate) higher-frequency strategies (on MES/ES/NQ/VX futures).
I'm thinking of writing some follow-up posts on how to reason about ML-driven strategies in an intraday setting. Thanks to low-cost brokerages, there's a lot of alpha that can be captured by small league speculators such as myself.
Thank you so much for your kind words! Your comment made my day! :)
Thank you!
I used Excalidraw (https://excalidraw.com), and I highly recommend it! It gives me 'xkcd' vibes.