HN user

tsazan

39 karma

Business Analyst specializing in AI commerce and cost optimization. Creator of CommerceTXT (reducing token overhead in agentic retail). Founder of PropertyFinder.bg - leveraging data analysis for international real estate markets.

Posts4
Comments36
View on HN

True. But extracting that metadata requires parsing the full DOM. CommerceTXT is for efficient discovery. Scan inventory cheaply first, then commit to the transaction.

The mapping approach assumes the web is static. In reality, you're building a 'maintenance debt' machine. For every 1,000 stores, you need 1,000 AI-generated mappings that break whenever a dev changes a CSS class.

CommerceTXT isn't just about extraction; it's about contract-based delivery. We are moving from 'Guessing through Scraping' to 'Knowing through Protocol'. You're optimizing the process of scraping; we are eliminating the need for it.

Because you don't need to audit every single transaction.

Think of it like a cache. You use the commerce.txt for 99% of your agentic workflows because it’s 30% cheaper in tokens and 95% faster than parsing a 2MB HTML haystack.

You only 'bother' with the HTML for periodic spot-checks or when a high-value transaction requires absolute verification.

Without CommerceTXT, you are forced to pay the 'HTML tax' on every single interaction. With it, you get a high-speed fast lane for context, while keeping the HTML as a decentralized source of truth for when trust needs to be verified. It’s about moving the baseline from 'expensive and fragile' to 'efficient and auditable'.

You’ve identified the exact tension we are navigating.

I support platforms like Shopify and Wix because they empower 80% of independent merchants to exist online. But I oppose their move toward 'enterprise-only' data silos. When Shopify gates their catalog API for a few select partners, they aren't protecting the merchant. They are protecting their own rent-seeking position.

CommerceTXT is a way for a merchant on any platform to say: 'My data is mine, and I want it to be discoverable by any agent, not just the ones who paid the platform's entry fee'.

Regarding 'design smell': Every major shift in computing has required specialized protocols. We didn't use Gopher for the web, and we shouldn't use 2010-era REST APIs for 2025-era LLMs. Models have unique constraints-token costs and hallucination risks-that traditional APIs simply weren't built to handle.

We aren't building for the gatekeepers. We are building for the open commons.

A CSV is a dump of facts. CommerceTXT is a layer of intent and logic. If you give an AI a giant CSV of your whole inventory, you blow the context window before the conversation even starts. If you serve a CSV per product, you still pay for headers and commas without getting any behavioral control.

Our spec handles this via @SEMANTIC_LOGIC and @BRAND_VOICE. It’s about how the AI represents your brand, not just the raw numbers.

Regarding bs4: mapping HTML to a thousand different store layouts is exactly what we are trying to escape. That is the 'fragility tax'. We are proposing a deterministic fast-lane that bypasses the need for custom scrapers for every single store.

You don't want the AI to 'guess' your data. You want it to 'know' your data.

That would be a fantastic first implementation. Openship is exactly the kind of architecture CommerceTXT is built for. Integration is straightforward: it’s essentially just a new 'View' layer. Instead of rendering HTML, you render a .txt endpoint that maps your existing product DB to our fields. I'll head over to your repo and open an Issue to discuss how we can map Openfront's data to the spec. I'd be happy to guide the implementation myself. Let's get this moving!

JSON is lean for data exchange between machines. But in the LLM economy, the currency is tokens, not bytes. To an LLM tokenizer, every bracket and quote is a distinct cost. In our tests, this 'syntax tax' accounts for up to 30% of the payload. We chose a line-oriented format to minimize overhead and maximize the context window for actual commerce data.

Agreed. In a perfect world, they would. But I cannot merge PRs into Shopify's core. Waiting for trillion-dollar corporations to change their security models is a death sentence for a new protocol. We build for the infrastructure that exists today, not the one we wish for. When they open the gates, we will move. Until then, we live in the root.

That solves the Token Tax. It fails the Bandwidth Tax. To get that JSON-LD, you still download 2MB of HTML. You execute JS. You parse the DOM. You are buying a haystack to find a needle, then cleaning the needle. We propose serving just the needle. Furthermore, JSON-LD is strictly for facts. It cannot express @SEMANTIC_LOGIC. It lacks the instructions on how to sell.

I respect that orthodoxy. It is the bedrock that allows the Internet to function. But we are optimizing for different variables. You optimize for architectural purity on a timeline of decades. You protect the namespace from temporary corporate flaws. I optimize for utility on a timeline of now. I want the flower shop owner to be visible to AI today, even if their platform is rigid. We have different North Stars. That is okay. You guard the temple. I will help the merchants outside. No leg-gnawing required. Thank you for the perspective.

You are right. Standardization often drifts from reality. That is why we built Section 9: Cross-Verification. The HTML remains the audit layer. The Agent does not trust blindly. It spot-checks. If commerce.txt says $50 but the HTML says $100, the merchant gets a Trust Score penalty. We do not replace the ground truth. We cache it, and we audit the cache to ensure it matches.

That solves bandwidth. It fails on tokens. JSON syntax is heavy. Brackets and quotes consume context window. More importantly, Schema.org is a dictionary of facts. It lacks behavior. It defines what a product is, but not how to sell it. It has no concept of @SEMANTIC_LOGIC or @BRAND_VOICE. We need a format that carries both data and instructions efficiently. JSON-LD is too verbose and too static for that.

It definitely lowers the barrier. But relying on messy HTML as a defense against competitors is 'security through obscurity'. It does not stop them; it just costs you server CPU. The data is public. If you put it on the screen, a scraper can read it. CommerceTXT just ensures that the good bots (AI Agents bringing customers) get it efficiently, while you can still block the bad ones via WAF.

It targets Consumer Protection and Truth-in-Advertising laws globally. The 'compliance bit' is Price Transparency. If an AI quotes a price as 'final' but checkout adds hidden fees or tax, that is a deceptive practice. Our spec enforces fields like TaxIncluded and TaxNote. It instructs the Agent to disclose whether the price is net or gross. It prevents the AI from accidentally committing fraud via misleading omissions.

The CC0 license is not a bug. It is a feature. If you fork this and build a standard that helps merchants better, the mission succeeds. I will be the first to applaud. As for "We": It is an invitation, not a pretension. A standard cannot be a solo act. I am bootstrapping the working group. You are welcome to join it, disagreements and all.

You do not download the haystack. You traverse it. The architecture is fractal. The agent reads the Root. If the user wants "Headphones", it follows that specific link. It ignores the rest. It is lazy loading for context. Do not mirror your DB manually. For real stores, generate the files dynamically. It is a view layer, just like HTML or sitemap.xml. Real-time? Yes. Since it is a dynamic response, it reflects the DB state instantly. Cache-Control headers handle the freshness.

There is no central authority. The Trust Score is a conceptual framework, not a shared database. Each AI platform (OpenAI, Anthropic, Google) builds its own model. They retain full discretion. Agents do not talk to each other. They talk to users. If a score is low, the agent warns the user. It adds caveats or drops the recommendation. It does not broadcast to other bots.

I appreciate the detailed feedback (and the edits).

You are technically correct regarding IETF norms.

But you say: "Wix and Shopify have zero bearing on the standardization of the Web."

I fundamentally disagree. The Web is not just a namespace for engineers; it is an economy for millions of small businesses. If a standard is technically "pure" but unusable by 80% of merchants on hosted platforms, it fails the Web.

However, to respect the namespace: We will mandate checking /.well-known/commerce.txt first.

But we will keep the root location as a fallback. We prioritize accessibility for the "aspiring" shop owner over strict purity for the standards writer.

Schema.org is the dictionary for facts.

We map strictly to Schema.org for all transactional data (Price, Inventory, Policies). This ensures legal interoperability.

But Schema.org describes what a product is, not how to sell it.

So we extend it. We added directives like @SEMANTIC_LOGIC for agent behavior. We combine standard definitions for safety with new extensions for capability.