<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Posts on tosd - Theory of Small Decisions</title><link>https://tosd.tech/en/posts/</link><description>Recent content in Posts on tosd - Theory of Small Decisions</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Fri, 11 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://tosd.tech/en/posts/index.xml" rel="self" type="application/rss+xml"/><item><title>Gold Is a Location, Not a Guarantee</title><link>https://tosd.tech/en/posts/gold-is-not-a-guarantee/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0000</pubDate><guid>https://tosd.tech/en/posts/gold-is-not-a-guarantee/</guid><description>&lt;p&gt;The medallion architecture divides a lakehouse into three layers, and anyone can read&#10;them without instruction: bronze is what arrived, silver is what&amp;rsquo;s been cleaned, gold is&#10;what you can use.&lt;/p&gt;&#10;&lt;p&gt;Then two teams query the same gold table and report different results. Or someone&#10;discovers midway through an outage that nobody knows who approves changes there.&lt;/p&gt;&#10;&lt;p&gt;A layer tells you where the data lives. It never tells you what it promises.&lt;/p&gt;</description></item><item><title>The Chat Always Answers, and That's the Problem</title><link>https://tosd.tech/en/posts/the-chat-always-answers/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0000</pubDate><guid>https://tosd.tech/en/posts/the-chat-always-answers/</guid><description>&lt;p&gt;A chat on top of your warehouse feels like you&amp;rsquo;ve finally crossed the finish line. Users&#10;ask in plain language, the model generates SQL, and the endless queue to your data team&#10;simply disappears.&lt;/p&gt;&#10;&lt;p&gt;Then someone asks how many active customers you had last quarter. They get a number back.&#10;They carry it into a meeting. Nobody can say what it counted as a customer, or as active,&#10;or whether it left out test accounts.&lt;/p&gt;</description></item><item><title>C4 for Data Architecture: The Wrong Container</title><link>https://tosd.tech/en/posts/c4-for-data-architecture/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://tosd.tech/en/posts/c4-for-data-architecture/</guid><description>&lt;p&gt;Adopting C4 is sold as the end of diagram chaos. Four levels, one audience each, and&#10;everybody finally draws the same thing.&lt;/p&gt;&#10;&lt;p&gt;On a data platform it half works. The diagram gets drawn, the review approves it, and six&#10;months later nobody opens it to decide anything.&lt;/p&gt;&#10;&lt;p&gt;The level of detail was never the problem.&lt;/p&gt;&#10;&lt;h2 id="every-data-platform-diagram-looks-the-same"&gt;Every data platform diagram looks the same&lt;/h2&gt;&#10;&lt;p&gt;Draw the container level for three platforms at three different companies and you get&#10;three versions of one picture: an ingestion layer, object storage, a processing engine,&#10;an analytical warehouse, an orchestrator and a BI tool.&lt;/p&gt;</description></item><item><title>An Ontology for the Semantic Layer</title><link>https://tosd.tech/en/posts/ontology-for-semantic-layers/</link><pubDate>Wed, 02 Sep 2026 00:00:00 +0000</pubDate><guid>https://tosd.tech/en/posts/ontology-for-semantic-layers/</guid><description>&lt;p&gt;Semantic layers are usually sold as a place to put metric definitions. Write &lt;code&gt;revenue&lt;/code&gt;&#10;once, expose it everywhere, stop arguing in meetings.&lt;/p&gt;&#10;&lt;p&gt;Then two teams report different revenue anyway, both technically correct, and the&#10;argument moves from the dashboard into the definition file.&lt;/p&gt;&#10;&lt;p&gt;The definitions were never the hard part.&lt;/p&gt;&#10;&lt;h2 id="metrics-are-the-visible-layer-not-the-load-bearing-one"&gt;Metrics are the visible layer, not the load-bearing one&lt;/h2&gt;&#10;&lt;p&gt;A metric is an aggregation over a set of things. Change what counts as one of those&#10;things and the number changes, without the formula changing at all.&lt;/p&gt;</description></item><item><title>Batch or Real Time: A Tradeoff Usually Framed Wrong</title><link>https://tosd.tech/en/posts/batch-or-real-time/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://tosd.tech/en/posts/batch-or-real-time/</guid><description>&lt;p&gt;The question arrives as a technology question. Should this integration be batch or&#10;streaming? And it gets answered like one, by comparing engines, latencies and&#10;throughput numbers.&lt;/p&gt;&#10;&lt;p&gt;That framing is the problem. Latency is not the requirement. It is a property you buy,&#10;and the thing worth arguing about is how much of it you actually need.&lt;/p&gt;&#10;&lt;h2 id="start-from-the-decision-not-the-pipeline"&gt;Start from the decision, not the pipeline&lt;/h2&gt;&#10;&lt;p&gt;Data exists so somebody, or something, can act on it. So the useful question is not&#10;&amp;ldquo;how fresh can we make this?&amp;rdquo; but &lt;strong&gt;&amp;ldquo;how quickly does the action it triggers actually&#10;happen?&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>