<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://sahajgarg.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://sahajgarg.github.io/" rel="alternate" type="text/html" /><updated>2026-03-19T23:25:13+00:00</updated><id>https://sahajgarg.github.io/feed.xml</id><title type="html">Sahaj Garg</title><subtitle>Sahaj Garg is building a more natural way to interact with technology, and  reflecting along the journey.</subtitle><author><name>Sahaj Garg</name></author><entry><title type="html">The Displacement of Cognitive Labor and What Comes After</title><link href="https://sahajgarg.github.io/blog/cognitive-labor/" rel="alternate" type="text/html" title="The Displacement of Cognitive Labor and What Comes After" /><published>2026-03-18T07:00:00+00:00</published><updated>2026-03-18T07:00:00+00:00</updated><id>https://sahajgarg.github.io/blog/cognitive-labor</id><content type="html" xml:base="https://sahajgarg.github.io/blog/cognitive-labor/"><![CDATA[<h2 id="i-the-threshold-has-been-crossed">I. The Threshold Has Been Crossed</h2>

<p>We are past the point where the question is whether artificial intelligence will exceed human capability across most cognitive domains. It already has. The remaining question is not <em>if</em> but <em>when</em> the full implications arrive, and the “when” is measured in months and years, not decades.</p>

<p>I want to start with personal experience, because the abstract version of this argument is less convincing than what it feels like to live through the shift.</p>

<p>I graduated in the top five of my class at Stanford. My entire identity, for most of my life, has been built on intelligence: the ability to reason through complex problems, to synthesize information, to produce meaningful work. That identity is now obsolete in the way I previously understood it.</p>

<p>Here is what that looks like in practice. The concrete productivity shift is staggering. In the past week alone, I have on five separate occasions produced output that I would have previously estimated at four weeks of hands-on engineering work. Each took roughly 45 minutes with a single well-crafted, well-directed prompt in Claude Code. So long as I can frame what I want clearly, the tool handles the rest, with the exception of edge cases where the system’s output meets reality and needs manual adjustment. That ratio (four weeks to 45 minutes) is not a marginal improvement. It’s a change in kind.</p>

<p>But the shift goes beyond raw productivity. On a daily basis, when I have a question or thought I want to explore, my first reaction is to go to an AI tool instead of talking to another person. This is a fundamental shift from before, where I would tend to brainstorm with human partners before AI partners. I now use AI as my primary partner for emotional reasoning, intellectual exploration, and strategic thinking. The places where I still go to humans are specific and narrow: I talk to my co-founder or my coach, because they will sometimes surface a way of thinking about a problem that’s different from how I’m approaching it, without me having to ask.</p>

<p>It also changes how I reason, write, and construct arguments. This essay is an example. It was produced through an eight-hour conversation between me and Claude, where I used Wispr Flow to dump my thoughts for five minutes at a time, learned from how Claude critiqued them, poked holes in Claude’s arguments, and went back and forth. This is draft 41. The AI didn’t hand me a polished essay. What actually happened was a process of iterative refinement: I’d voice a rough idea, Claude would push back on the weak parts, I’d disagree with some of its pushback and accept other parts, and each cycle produced a tighter version of the argument. It took those eight hours for me to realize all the ways in which my initial framing was wrong and to arrive at a significantly more rigorous understanding of what I’d been thinking. That kind of interaction was not possible until very recently.</p>

<p>This is a remarkable change. I now view my intelligence as a commodity on tap. The primary skills I have left are my ability to synthesize different AI-generated perspectives into a combined whole that’s better than what the AI produces on its own, my ability to judge what I like and don’t like, and my ability to identify what I care about. Taste, direction, synthesis. Not raw cognitive horsepower. For someone whose whole identity was predicated on that horsepower, this is a significant transformation, and it’s one that millions of knowledge workers will go through in the next few years whether they’re ready for it or not.</p>

<figure class=""><img src="/assets/images/ai-abundance/Intelligence2.png" alt="Wait But Why&#39;s AI Intelligence over Time chart." /><figcaption>
      Tim Urban drew this in 2015. We’re somewhere just short of Einstein now. What happens when we move slightly to the right in four to eight weeks?

    </figcaption></figure>

<h2 id="ii-harnessing-intelligence-and-the-impacts-on-labor">II. Harnessing Intelligence and the Impacts on Labor</h2>

<p>For years, the open question in AI was whether scaling laws would hold: whether you could keep increasing compute, data, and parameters and keep getting predictable improvements in intelligence. That question is settled. The scaling laws held out long enough. The raw capability is here.</p>

<p>But raw capability is not the same as structured impact. AI today is roughly where human intelligence would be without the organizational structures we’ve built to harness it. A brilliant individual can do remarkable things. A brilliant individual leading a well-designed corporation, with teams and hierarchies and processes, can reshape an industry. AI has the brilliance. What it’s still developing is the organizational scaffolding. And because what remains is an engineering problem (how to harness, structure, and deploy intelligence effectively) rather than a research problem (whether the intelligence will materialize at all), the timeline is a matter of execution speed, not scientific discovery.</p>

<p>That scaffolding is arriving fast. The inflection point, for me, was watching Claude Code begin to orchestrate sub-agents effectively: assembling teams of AI workers to tackle different dimensions of a complex task in parallel, then synthesizing the results. This is a direct analogue to how human organizations function. You break a problem into pieces, assign them to capable actors, and integrate their output. The fact that AI can now do this, imperfectly but recognizably, means the path to fully autonomous knowledge work is a matter of engineering iteration, not fundamental research. That mental model scales to arbitrary levels of abstraction: agents coordinating sub-agents coordinating sub-sub-agents, mirroring the way a CEO directs VPs who direct managers who direct individual contributors. The innovation of making this work fluidly at scale hasn’t fully arrived, but it’s very clear that it will.</p>

<figure class=""><img src="/assets/images/ai-abundance/org_vs_agent_hierarchy_v2.svg" alt="Parallel structures: human org chart and AI agent hierarchy." /><figcaption>
      Same organizational logic, different substrate.

    </figcaption></figure>

<p>The other recent shift is the emergence of working computer-use agents. Perplexity’s computer-use tool more or less works: an AI system that can operate a computer the way a human does, launching applications, interpreting visual output, navigating interfaces, and iterating based on results. This closes a critical gap. Previously, AI could think but couldn’t act in digital environments without human intermediation. That barrier is falling.</p>

<p>Consider a concrete example. Say you need to hire Android engineers. Previously, you’d delegate to a recruiter or spend days doing it yourself: researching which companies have strong Android teams, finding individual engineers, compiling contact information, writing personalized outreach, sending connection requests one by one. A computer-use agent can do all of this autonomously. It compiles a list of the 20 most prominent Android companies, searches LinkedIn for the top engineers at each, puts them into a tabulated spreadsheet, drafts a personalized outbound message to each person, and controls LinkedIn to actually send those messages as connection requests. That’s not a chat interaction. It’s an agent operating a computer the way a human recruiter would, except it takes minutes instead of days. This is the kind of task that doesn’t get solved by talking to an AI in a text box. It gets solved by AI being able to use the same tools we use, in the same environments we use them. The ingredients for full automation of knowledge work are now all present.</p>

<p>It helps to think about human labor as falling into two broad categories. The first is cognitive labor: knowledge work and emotional work. Analysis, writing, coding, design, legal reasoning, financial modeling, medical diagnosis, therapy, coaching, project management. This is the category being automated right now. The second is physical labor: construction, manufacturing, agriculture, logistics, plumbing, electrical work, maintenance. This requires robotics and physical automation, which follows a different timeline.</p>

<h2 id="iii-physical-intelligence-what-is-to-come">III. Physical Intelligence: What Is to Come</h2>

<p>People often joke that you can’t use nine people to make a baby in one month. The implied lesson is that some things are irreducibly serial and can’t be parallelized. This is true, but it misses what AI actually changes about physical automation timelines.</p>

<p>The bottleneck in developing robotics and physical intelligence has always been the cognitive labor of R&amp;D: designing systems, running experiments, analyzing results, iterating on designs. That cognitive labor is precisely what’s being automated now. When AI can run massively parallel experimentation strategies for robotics development, the limiting factor is no longer human researchers working sequentially. The limiting factor becomes the physical production cycle itself: building prototypes, testing them in the real world, observing how they break. You still can’t skip the serial contact with physical reality, but you can compress everything around it.</p>

<p>This is a fundamental shift. Development cycles that previously took months of human cognitive work sandwiched between physical tests now take days of AI cognitive work sandwiched between the same physical tests. The serial part (physical production and testing) stays the same length, but the parallel part (design, analysis, iteration) compresses by an order of magnitude or more. And if design and testing become nearly free, you get to test many designs in parallel, limited only by budget for physical prototypes rather than by human cognitive bandwidth. The result is that improvements in robotics and physical autonomy will arrive far faster than most people’s intuitions predict.</p>

<p>Robotics today is already decent, and the compounding dynamic is powerful. Every advance in AI capability accelerates every subsequent advance, including in the physical domain. If cognitive labor automation is largely complete within three to five years, and that compressed R&amp;D cycle is applied to robotics, physical labor automation could follow within five to ten years rather than the twenty or thirty that a linear extrapolation would suggest.</p>

<p>This matters enormously for the transition. If physical labor automation arrives on a ten-year timeline, it means we cannot treat physical labor as a permanent refuge for displaced knowledge workers. Unlike the Industrial Revolution, where one type of human labor (manual agricultural work) was replaced but another type (factory and later knowledge work) grew to absorb the displaced population, this transition has no obvious next category of labor for humans to move into. All three categories (knowledge work, emotional work, physical work) are converging toward automation within a single generation.</p>

<figure class=""><img src="/assets/images/ai-abundance/labor_automation_timeline_v4.svg" alt="Labor displacement curves showing knowledge and emotional work automated before physical work." /><figcaption>
      The gap between cognitive automation and physical automation is the acute transition period.

    </figcaption></figure>

<h2 id="iv-abundance-and-continued-scarcity">IV. Abundance and Continued Scarcity</h2>

<p>Not everything becomes abundant in the same way or on the same timeline, and the split between what does and what doesn’t is one of the most important features of the post-economic world to understand clearly.</p>

<p>A lot of people in Silicon Valley are assuming that everything will become abundant. Elon Musk talks about “universal high income” as though the problem is simply distributing the surplus. But that framing obscures a critical distinction: some things will approach zero marginal cost, and some things will not. Figuring out how to manage a world where some goods are effectively free and others remain scarce (or become more valuable) is one of the defining challenges ahead.</p>

<p><strong>Things that will become radically abundant</strong> (approaching zero or very low marginal cost):</p>

<ul>
  <li>Information synthesis, analysis, and generation. Any question you can ask will have a better answer than any individual human expert could provide, instantly, for free.</li>
  <li>Legal document drafting, contract review, and regulatory compliance. The entire apparatus of routine legal work becomes automated.</li>
  <li>Software generation for standard applications. Custom software becomes as cheap as asking for it.</li>
  <li>Translation and transcription, in real-time, across any language pair.</li>
  <li>Basic medical diagnosis and triage. AI systems that outperform most human diagnosticians, available to everyone.</li>
  <li>Financial modeling, analysis, and planning. Personalized financial advice that today costs thousands of dollars per hour becomes free.</li>
  <li>Education content and tutoring. Personalized instruction tailored to each learner’s pace, style, and level, infinitely patient, available 24/7.</li>
  <li>Design variations and iteration. Producing a hundred design options becomes as easy as producing one.</li>
  <li>Customer support and routine communication. Any interaction that follows patterns gets handled by AI.</li>
  <li>Physical goods produced by automated systems. Once robotics matures, manufactured goods, constructed buildings, and produced food approach the same zero-marginal-cost dynamic that digital goods reach first.</li>
</ul>

<p>All of these currently require paying enormous amounts: a lawyer charges hundreds per hour, a software engineer costs $150,000+ per year, a financial advisor takes a percentage of your assets. The cost of equivalently intelligent AI models is dropping rapidly. Even at today’s prices, AI is far cheaper than paying a human for the same cognitive output. As costs continue to fall, these services will become truly abundant.</p>

<p><strong>Things that will NOT become abundant</strong> (and may become more valuable):</p>

<ul>
  <li>Physical presence and embodied experience. Your Taylor Swift concert. The feeling of being in a specific place at a specific time with other humans.</li>
  <li>Physical scarcity. Land, nature, raw materials, pristine environments. There is one beachfront lot in that spot, and no amount of AI changes that.</li>
  <li>Health and physical excellence. Athletic achievement, physical skill, the experience of a body that works well. These require a human body and can’t be outsourced.</li>
  <li>Craft and artisanal production. The “human-made” premium. A hand-carved table carries value precisely because a human made it and you know that. Homemade baked goods carry value precisely because you made them yourself.</li>
  <li>Taste, curation, and judgment. Knowing what to ask for, knowing what’s good, having an opinion that’s genuinely yours.</li>
  <li>Social proof and status signaling. Being known, being vetted, having a reputation that other humans vouch for.</li>
  <li>Novel lived experience. Travel, adventure, relationships. The things that require you to actually be present in your life.</li>
  <li>Spiritual and contemplative goods. Retreat experiences, contemplative communities, the fruits of sustained practice.</li>
  <li>Community and belonging. Clubs, movements, shared purpose. The experience of being part of something with other people.</li>
  <li>Political representation and advocacy. Someone fighting for your interests, accountable to you.</li>
  <li>Risk-bearing and skin-in-the-game decisions. The person who signs the liability waiver. The human who takes responsibility when things go wrong.</li>
</ul>

<p>Economics doesn’t break down. But a good chunk of capitalism might. Capitalism operates on the basis of prices aggregating information to coordinate production and consumption. When the cost of producing many goods and services approaches zero, that price mechanism stops functioning for those goods. There’s no price signal to send when the marginal cost of one more unit of legal advice, or software, or medical diagnosis is effectively nothing. It’s an open question which parts of capitalism break down and which adapt.</p>

<p>Consider how we already handle this for air. We think of air as free, but the Clean Air Act created a market by placing a high price on pollution, and people pay a premium to live in places with cleaner air. There is an economy around air; it just looks nothing like the economy for scarce goods. As cognitive goods and services approach zero marginal cost, our economic systems will need to adapt in similar ways, developing new frameworks that look as different from today’s markets as emissions trading looks from buying groceries. The transition we’re heading into is the process of moving everything in the first list from “scarce” toward something closer to “air,” while the second list remains the terrain where competition, allocation, and human striving still live. What’s clear is that the frameworks we have today were not designed for a world where scarcity applies to far fewer things than before, and in different ways. The implications of this shift will be one of the most complex questions of the coming decades.</p>

<h2 id="v-the-transition">V. The Transition</h2>

<p>The endpoint of abundance may be desirable, at least to many people born into it. The transition to it will likely be the most turbulent period in modern history, not because of AI catastrophe, but because of the speed at which existing social and economic structures will be disrupted.</p>

<h3 id="the-knowledge-work-cliff">The Knowledge Work Cliff</h3>

<p>The first and most immediate disruption is the automation of knowledge work. Within three to five years, the majority of cognitive jobs will be substantially automated. This doesn’t mean every knowledge worker loses their job overnight. It means the number of humans needed to produce a given unit of cognitive output drops by an order of magnitude. One person with AI does what twenty did before. The math is simple: most knowledge workers become redundant.</p>

<p>This hits the upper-middle class hardest, which is counterintuitive. The truly wealthy have diversified assets and capital reserves. Working-class people doing physical labor still have scarce skills. It’s the $80,000-to-$400,000 household (the lawyer, the software engineer, the financial analyst, the radiologist) that gets caught. Their skills are the most directly substitutable by AI, and their lifestyle is built on continuous high income rather than accumulated capital. They lose their income, and the value of their primary asset (typically a home in an expensive knowledge-work city) declines as demand in those cities contracts.</p>

<p>The instinctive policy response is retraining: move displaced knowledge workers into physical labor, which still requires humans. This mostly won’t work. There won’t suddenly be demand for twice the number of people in physical labor industries just because half of knowledge work is eliminated. The skills gap takes years to close. And as I argued in the previous section, physical labor is also on the path to automation, just on a delayed timeline. There is no permanent refuge.</p>

<h3 id="identity-not-just-income">Identity, Not Just Income</h3>

<p>The economic disruption is severe, but the identity crisis may be worse. We have a direct historical parallel: the deindustrialization of the American Rust Belt.</p>

<p>When manufacturing left cities like Detroit, Youngstown, and Flint, the economic damage was obvious. But the deeper wound was to identity. These were communities where being a steelworker or autoworker wasn’t just a job; it was who you were. It defined your family’s place in the social order, your sense of contribution, your self-respect. As Case and Deaton wrote in <em>Deaths of Despair</em>: “Jobs are not just the source of money; they are the basis for the rituals, customs, and routines of working-class life. Destroy work and, in the end, working-class life cannot survive.” The loss of that identity, even more than the loss of income, drove the epidemic of depression, substance abuse, and what they termed “deaths of despair.”</p>

<p>The coming displacement of knowledge workers has the potential to be that dynamic at a much larger scale. A software engineer’s identity is as tied to their cognitive ability as a steelworker’s was to their trade. “I’m smart, I solve hard problems, I build things” is not just a job description. It’s a self-concept. When AI can solve harder problems and build things faster, that self-concept shatters.</p>

<p>By default, I suspect that mass unemployment of knowledge workers, simultaneous with the dissolution of a core part of their identity, will lead to widespread depression and despair.</p>

<h3 id="political-stability">Political Stability</h3>

<p>I want to be clear that what follows is speculative. Nobody knows exactly how this plays out politically. But the historical patterns are concerning enough to take seriously.</p>

<p>Rapid displacement of economically powerful classes has historically been destabilizing. The French Revolution wasn’t driven by the poorest; it was driven by the bourgeoisie who felt their status didn’t match their economic reality. The Arab Spring was fueled in large part by highly educated young people who had invested in university education and found themselves with no economic opportunity. In Jordan, 53% of unemployed youth held university degrees. Across the region, the message had been: get educated, and opportunity will follow. When that promise proved hollow, the sense of betrayal was more radicalizing than poverty alone.</p>

<p>The parallels to a potential mass displacement of knowledge workers are hard to ignore. These are the most educated, most politically engaged, most socially connected demographic in society. Many have taken on significant student debt for degrees that promised professional careers. If those careers evaporate, the combination of economic loss, identity dissolution, and a sense of having been deceived by the institutions they trusted could be a potent formula for instability. Whether that instability takes the form of productive political organizing or something more destructive is genuinely uncertain.</p>

<p>The AI infrastructure companies are likely candidates for some redistribution during this period, driven more by self-interest than altruism. They need a functioning society to operate in, and the most likely form this takes is making AI services extremely cheap or free, which reduces the cost of living even as incomes fall. But this creates its own uncomfortable dynamic: a large population dependent on the goodwill of a handful of companies, without democratic accountability or structural guarantees.</p>

<h3 id="governance-and-geopolitics-during-the-transition">Governance and Geopolitics During the Transition</h3>

<p>These next points are also speculative, but they follow logically from the dynamics above.</p>

<p>AI makes produced goods abundant, but land, location, and certain physical resources remain zero-sum. The property rights and governance frameworks we have today were built for a world of broad scarcity. When scarcity narrows to a small set of rivalrous assets, the rules for who gets what and how those decisions are made will need to be fundamentally reimagined. What does property ownership mean when construction is free but location is priceless? How do you govern land use when the economic logic that previously drove zoning and development no longer applies? I don’t have confident answers here, and I’m not sure anyone does yet.</p>

<p>Similarly, if the United States, or more precisely a handful of US-based companies, achieves something approaching post-economic abundance domestically while the rest of the world does not, the resulting asymmetry could be the largest in human history. Not inequality of wealth, but inequality of productive capacity itself. The frameworks we use to understand international relations are built on economic interdependence, scarcity, and trade. What happens to those frameworks when some nations no longer need others for production is genuinely unclear. It could lead to a more stable world (abundant nations have less reason for conflict) or a far less stable one (nations facing permanent irrelevance don’t accept it quietly). The range of outcomes is wide, and I suspect the specifics will depend heavily on decisions made in the next decade.</p>

<h2 id="vi-post-abundance-human-nature-doesnt-change">VI. Post-Abundance: Human Nature Doesn’t Change</h2>

<p>It’s important to separate the challenges of the transition from the challenges of what comes after. The transition is about displacement: jobs disappearing, identities breaking, institutions failing to keep pace. The post-abundance world is a different problem entirely. It’s about fundamental human nature meeting a set of conditions that human nature has never encountered before.</p>

<p>The technology changes. Human nature does not. The drives that define us (status-seeking, hierarchy, competition, the desire to be valued relative to others) are not products of capitalism or knowledge work. They predate both. These dynamics existed before knowledge work was important to society, and they will continue after knowledge work is automated. What changes is the terrain on which these drives play out, and the social contracts that channel them.</p>

<h3 id="for-those-not-yet-born">For Those Not Yet Born</h3>

<p>There is an important asymmetry between people alive today and those who will be born into a post-economic world. For today’s adults, the transition involves a painful identity transformation: letting go of professional identity, status hierarchies, and assumptions about what constitutes a valuable life.</p>

<p>For children born into this world, abundance will simply be the water they swim in. We may be the last generation that has to work for a living. Our children will likely grow up in a world where having a career is a choice, not a necessity. They won’t experience abundance as remarkable any more than we experience indoor plumbing as miraculous (and it’s worth remembering that in 1940, 45% of American homes still lacked indoor plumbing; within living memory, this was a luxury, not a given). Their sense of identity and status will not be anchored to career or professional accomplishment, because those concepts will carry no more weight than “being good at churning butter” carries for us today. These attachments never form in the first place. Their developmental challenges will be different. Not the grief of losing a role, but the challenge of constructing meaning without the scaffolding of material necessity.</p>

<p>When people raise this concern, a common response is that humans fundamentally need work for meaning, and that a world without work is therefore a world without meaning. I think this claim is weaker than people realize. It’s really a claim about a specific cultural arrangement (Protestant work ethic, American meritocratic ideology) rather than a deep truth about human nature. Stay-at-home parents find deep meaning in raising children. Retirees find it in gardening, volunteering, grandchildren. Religious contemplatives find it in practice and community. Many cultures, particularly in Southern Europe and the non-Western world, have never placed work at the center of identity the way American culture has.</p>

<p>What people actually seem to need is not work specifically but four things that work happens to provide: agency (the sense that you’re making choices that matter), contribution (the sense that you’re valued by others), mastery (the sense that you’re getting better at something), and connection (belonging to something larger than yourself). Work can provide all four, but it’s not the only thing that can. The challenge is building new structures that provide these when work no longer does.</p>

<p>Gen Z is an interesting case here. A generation that already doesn’t derive primary identity from work is in some ways better positioned for this transition. But they’re also potentially worse off, because the existing social infrastructure (healthcare tied to employment, housing tied to income, social status tied to job title) hasn’t caught up to their values.</p>

<h3 id="status-and-competition-after-work">Status and Competition After Work</h3>

<p>Human status-seeking doesn’t disappear when material needs are met. It redirects. In a world where AI can produce any material good, conspicuous consumption becomes meaningless as a status signal. So status migrates to whatever remains scarce: physical excellence, social and relational capital, creative originality, governance influence, spiritual attainment. The likely candidates map closely to the “things that will not become abundant” list from Section IV.</p>

<p>Which of these dominate matters enormously. The pessimistic outcome is that status collapses into the most visible and competitive dimensions: physical appearance, social following, zero-sum games. The optimistic outcome is a restructuring around depth: wisdom, psychological development, quality of relationships, contribution to community. The mechanisms change. The underlying drive doesn’t.</p>

<p>The social contract of our society for how we channel these drives will have to be fundamentally reimagined. The current contract says: compete through economic productivity, and status follows. That contract breaks when economic productivity is automated. The replacement contract doesn’t exist yet, and building it is one of the central challenges of the post-abundance world.</p>

<h2 id="vii-open-questions">VII. Open Questions</h2>

<p>The trajectory I’ve described raises a set of questions that I don’t think anyone currently has good answers to. I’m going to explore these in future essays, but I want to name them here because I think they’re the most important questions of the next two decades. If you’ve thought deeply about any of these and believe you have a path forward, I’d like to hear from you.</p>

<p><strong>On the transition:</strong></p>

<ul>
  <li>What happens to the displaced class of knowledge workers? Not in the abstract, but concretely: how do tens of millions of highly educated, previously high-status people navigate the simultaneous loss of income, identity, and social position?</li>
  <li>What are the social and political implications of this displacement? Is the default path one that leads to unionization, uprising, revolution, or something else?</li>
  <li>Could these dynamics lead to targeted violence against AI companies and AI infrastructure?</li>
  <li>Could they lead to actual political revolution or regime change?</li>
  <li>How might we politically manage this transition period to allow for stability? What institutions, policies, or agreements could prevent the displacement from becoming destabilizing?</li>
  <li>By default, I suspect that mass unemployment of knowledge workers, simultaneous with the dissolution of a core part of their identity, will lead to widespread depression and despair. How do we handle that at scale?</li>
  <li>What are the geopolitical implications of one or two nations developing this technology before other nations? What happens to the international order when economic interdependence dissolves?</li>
</ul>

<p><strong>On post-abundance:</strong></p>

<ul>
  <li>Is the default path for large groups of people to end up in something like Brave New World or Ready Player One: a society optimized for comfort and stimulation rather than growth and meaning? Is there a way to avoid that default path?</li>
  <li>How do we preserve and create a world with human flourishing where labor is largely displaced?</li>
  <li>How does governance and economics work when some things are scarce and other things are not? What property rights frameworks make sense in a world where production is free but location is priceless?</li>
</ul>

<p>These are not rhetorical questions. They are the defining challenges of the next generation, and the window to shape the answers is now.</p>

<hr />

<p><em>Acknowledgements: I want to thank my friend and colleague Michael Hochberg, who was instrumental in crafting this essay by pushing back on many arguments I had absorbed from the closed Silicon Valley dialogue that can be a self-reinforcing bubble. His <a href="https://longwalls.substack.com/p/the-ai-grunderzeit">article on the AI Gründerzeit</a> is an excellent perspective worth reading.</em></p>]]></content><author><name>Sahaj Garg</name></author><category term="blog" /><category term="AI" /><summary type="html"><![CDATA[I. The Threshold Has Been Crossed]]></summary></entry><entry><title type="html">Recommended Reads</title><link href="https://sahajgarg.github.io/blog/recommended-reads/" rel="alternate" type="text/html" title="Recommended Reads" /><published>2026-02-26T21:52:00+00:00</published><updated>2026-02-26T21:52:00+00:00</updated><id>https://sahajgarg.github.io/blog/recommended-reads</id><content type="html" xml:base="https://sahajgarg.github.io/blog/recommended-reads/"><![CDATA[<p>Below I share books that have shaped me in different ways:</p>
<ul>
  <li>Most influential</li>
  <li>Most moving</li>
  <li>Most delightful and fun to read</li>
  <li>Most useful</li>
</ul>

<h2 id="most-influential">Most Influential</h2>

<p><strong>Siddhartha by Hermann Hesse</strong></p>

<p>This is the single most influential novel I’ve read. It’s a two-hour delightful read that tells the story of a character named Siddhartha on a parallel journey to the Buddha on the path to enlightenment. It serves as a reminder that we must all craft our own journey through life. We must learn by being, not just reading.</p>

<p>There are so many themes packed into this short novel that it’s hard to know where to start. The one that strikes me most is the tension between experience (wisdom?) and knowledge — how there are severe limits to what can be thought or conveyed through words or ideas alone. Ironically, it was the experience of walking the Camino de Santiago — a weeks-long pilgrimage across Spain — that led me to experientially understand the importance of experience over knowledge. I might have read sections about how there were things that could not be learned from a teacher, but I certainly did not appreciate or understand those bits until I’d lived them. The novel is also a powerful reminder of our preoccupation over trivialities and our fictionalization of concerns — especially at work, echoing characters who lose themselves in commerce and routine. And finally, it’s full of departures — a son leaving his father, lovers parting ways, a parent watching a child walk away — that tap into something universal about growing up: the guilt and grief of leaving the people who love you most, knowing you’ll never really come back in the same way. What a gold mine of themes and ideas!</p>

<p><strong>Brave New World by Aldous Huxley</strong></p>

<p>This book was influential to me due to the timing — it served as a reminder to challenge all assumptions about the world around us. The things that we learn are conditioned to us from birth by our education system, and very little of our thinking is truly independent. <em>Brave New World</em> served as a wake-up call for questioning assumptions I have held most my life.</p>

<p><strong>Sex, Ecology, and Spirituality by Ken Wilber</strong></p>

<p>The book outlines Wilber’s integral theory — essentially a universal theory of philosophy, science, consciousness, and culture. It does so by characterizing the nature of the world as being composed of holarchies: in essence a hierarchy where each level in the hierarchy is both a whole, and a part of the next level. Growth in each hierarchy involves transcending the current level by negating it, but then preserving and integrating it. This mental model helps explain inner consciousness development, the physical material world, culture, and observable cognitive development — and the levels of these hierarchies are related across all dimensions. Personally, what I take away from this book is the simultaneous pursuit of Ascent and Descent — of striving towards deeper and deeper transrational consciousness while embracing a love for the world as it is. More broadly, I take away the notion of integration: I can learn from a diversity of perspectives, most individual perspectives are incomplete, and integrating them is one of the most profound tools for developing deeper consciousness. Overall, a book I’d highly recommend to anyone looking to expand their way of thinking and being in the world, but it’s very dense.</p>

<h2 id="most-moving">Most Moving</h2>

<p><strong>When Breath Becomes Air by Paul Kalanithi</strong></p>

<p>The most emotionally moving book I have ever read — the story of a neurosurgeon who passes from brain cancer. I spent four hours more or less continuously crying while reading the book in a single sitting. Kalanithi paints a beautiful and devastating picture of mortality.</p>

<p><strong>Steppenwolf by Hermann Hesse</strong></p>

<p>I see myself in this novel; torn between a multitude of personalities and identities that I apply in different situations, with a split mindset between generations. It’s a gripping psychological fantasy novel that reflects tensions in Hesse’s underlying identity as half-man and half-wolf.</p>

<p><strong>Being Mortal by Atul Gawande</strong></p>

<p>Most of us don’t think about mortality much on a day-to-day basis, but it’s one of the most defining parts of the human experience. This book prepares us for mortality and how modern medicine treats death.</p>

<p>Makes you see aging in a different light. Before reading this, I assumed that putting a parent in a nursing home was mostly a sign of neglect. This book completely changed that view — and opened my eyes to the kinds of new alternatives that can make the end of life more meaningful. Full of content and ideas that I had never heard or read about. I will take away many lessons from this book for the care of people around me who are important to me, and for myself (hopefully many years down the line). Extremely moving, brought me to tears many times.</p>

<h2 id="most-delightful-and-fun-to-read">Most Delightful and Fun to Read</h2>

<p><strong>The Three-Body Problem by Cixin Liu</strong></p>

<p>A brilliant sci-fi book, so fun to read. I put it down the first two times I picked it up because the prose is bland — but the story is gripping. The cultural commentary is insightful, and it exposed me to the idea of “new science,” i.e. scientific paradigms that we have no conception of today, that are beyond observation with our current tools, but that may contain truths about our universe. I’ve read very little that feels as creative and imaginative as this book.</p>

<p><strong>Don Quixote by Miguel de Cervantes</strong></p>

<p>The first novel ever written, and a very, very funny story. I enjoyed this as part of learning the history of the novel.</p>

<p>One of my favorite books, a delight to read. I never realized how funny a novel could be until reading <em>Don Quixote</em>. The novel pulled me back every week for another fifty pages, and never became tiresome. It brought me a reminder of literature for literature’s sake, as opposed to literature for ideas (although I certainly took away that meta-level idea). Its episodic nature made it particularly fun, like a sitcom in a book. It also serves as a reminder of the strangeness of humanity: it’s not just today that people’s minds are strange, but it’s always been so, and if anything we’re more moderate and tempered and uptight than ever.</p>

<p><strong>The Rosie Project by Graeme Simsion</strong></p>

<p>The book is a story of a person with Asperger’s and how he views the world. The portrait it paints is so vivid, and helps step into someone else’s shoes better than almost any other novel I’ve read.</p>

<p><strong>Brief Interviews with Hideous Men by David Foster Wallace</strong></p>

<p>This is a short story collection, and has the most riveting prose of anything I have ever read. The sheer brilliance of most of the short stories more than makes up for the ones that are maddeningly difficult to get through. The collection gets its power from its brutal honesty and its incisive character portraits — I’ve never seen depictions of people like this anywhere else. Wallace exposes the emotional perversions of his characters with an honesty that reminds you we all have hidden ugliness, and that we can do better to confront it (at least in ourselves). The anger can get a bit exhausting at times, but as a whole the collection is just brilliant. Besides the content, it’s stylistically phenomenal: the wordplay, the structural play, the meta-fiction on meta-fiction, the lens into how Wallace writes his stories in the first place (<em>Octet</em>), the footnotes, the omissions of responses in dialogue. I’ve never seen anything quite like it. Some moments that stood out: the man who counts a woman’s orgasms being just as hideous since people crave to please their partners too; the inability to communicate the feeling of depression; the recognition of self-absorption and that trapping the person listening. Honestly just so many vivid feelings that I’ve never seen captured quite this well anywhere else.</p>

<h2 id="most-useful">Most Useful</h2>

<p><strong>The Mom Test by Rob Fitzpatrick</strong></p>

<p>This book is a must if you’re starting a company or building products. It talks about how to get user feedback: if you ask someone about an idea you’re working on, they’re bound to say it’s good (like your mom would), just to make you happy. This book sets up fantastic frameworks for getting actual user feedback. Implementing it is much harder than it sounds!</p>

<p>This book feels like an application of the scientific method to customer conversations — something that’s natural for a skeptic. It’s also one of those books that people love to evangelize and then do exactly the opposite of — I’ve watched people ask leading questions and then declare “see, everyone loves what we’re building!” Some of the biggest takeaways:</p>

<ul>
  <li>If they haven’t already tried to solve their problem, it probably isn’t that big of a problem for them.</li>
  <li>Dig around at emotional signals — when you’re building technical products, it’s easy to focus only on the tech.</li>
  <li>“What are your big goals and focuses right now?” is a great question for understanding people’s priorities.</li>
  <li>Talk to the multiple people who could be failure points — you’re not usually selling to a single person.</li>
  <li>Don’t overdo the meetings — they’re about quickly learning what you need.</li>
</ul>

<p><strong>The Defining Decade by Meg Jay</strong></p>

<p>Jay writes about the 20s and how the decisions we make in our 20s impact the rest of our life. A lot of it could be summarized, but reading it fully helps the lessons internalize and helps lead a life with less regret.</p>

<p>A refreshing, quick read and a relatively standard self-help book. There are three major sections. The one on your career is interesting if you’re not already ambitious; in my case, it felt mostly like a reminder that the decision was in my hands about what I would end up doing. The one on relationships is probably the most broadly useful. It’s a good antidote to the commitment-averse tendencies you see in Gen Z and Gen Alpha — the instinct to keep all options open, to treat doubt as a dealbreaker rather than a normal part of being with someone. The final section on mind and body was probably the most useful to me, and to most ambitious people. The core idea — that twentysomethings with overactive amygdalae often want to change their feelings by changing their circumstances, quitting things that get messy rather than sitting with necessary imperfection — really resonated with me and was a strong reminder of how tempting it is to run from challenges instead of through them. The final remarks on fertility statistics were a good reminder of why it’s worth thinking about having kids earlier than the coastal elite default, and accounting for that in your planning.</p>

<p><strong>Flourish by Martin Seligman</strong></p>

<p>This book outlines a framework for human flourishing. Seligman breaks down the primary ingredients of a meaningful life into five elements — Positive emotion, Engagement, Relationships, Meaning, and Achievement (PERMA) — which helps clarify where to actually devote mental energy. Too much of our modern world is focused on achievement alone over personal fulfillment and connection with the people around us, and this framework is a useful corrective for rebalancing.</p>

<p><strong>Never Split the Difference by Chris Voss</strong></p>

<p>The best handbook for negotiation ever written! Brilliant, and full of actionable advice. Some of the takeaways that resonate with me most strongly right now include: using calibrated (“how”/open) questions to get the other person to work to solve your problems (“How am I supposed to pay that much?” e.g., or “Where can I find that money?”); saying no softly (“I appreciate that offer, but that won’t work for my budget.”); using specific numbers (this one still feels weird, but is likely effective despite that); preserving the decision maker’s agency, and understanding their non-monetary factors at play (often will be black swan factors — maybe the person is up for promotion?); labeling the other party’s fears/emotions/concerns as a way to abate them (“You may be worried that we are young founders.”), in particular pause after stating a label to allow them the space to respond; getting no answers over yes (no maintains agency, yes feels like something you say to get the conversation over with).</p>

<p><strong>No Bad Parts by Richard Schwartz</strong></p>

<p>Insightful (but not exceptional as a read). Mostly gave it five stars because I suspect it will have a deep lingering impact on me three months from now. It’s one of those kinds of books that steeps into my mind and will stay with me as a lens.</p>

<p>There’s a good chance that this book helps me discover the source of a handful of patterned behaviors that irritate me — I’ve solved most of my general stressors but there’s still a handful that linger and create nuisances for me.</p>

<p>The most revelatory part of the mental model to me was thinking of parts as: 1) in multitude — I’ve often thought of people as having a couple of parts at most, sort of like <em>Steppenwolf</em>, but never a multitude, 2) wholes, in and of themselves — each emotion reflects an identity that’s operating in the driver’s seat, and 3) as third-party identities, where I can ask what is “it” angry about instead of what am “I” angry about, which changes the mental model.</p>

<p>Excited to see what tangible impacts this book makes on my life.</p>

<p><strong>Scaling People by Claire Hughes Johnson</strong></p>

<p>An excellent book that I read at the perfect time for myself. This book is full of tactical advice for how to scale organizations, improve interpersonal dynamics/relationships, and uplevel yourself and others. It’s one of the only resources I’ve found that has actual tactical, operational advice for people management. If you’re a very intentional leader, you might learn many of these things on your own over time, but I was encountering many of them for the first time as my org hit ~20 people while I was reading the book. Most other books on people management are philosophies/platitudes/generalities, not tactics that are tangible and implementable. If you’re looking for tactical advice for improving people management, this is the place to go!</p>

<p><strong>Outlive by Peter Attia</strong></p>

<p><em>Outlive</em> was way better than I expected. I really wanted to dislike it (especially given I dislike the people who appear to be Attia-cult-followers all throughout Silicon Valley), but it was a fantastic read.</p>

<p>What I appreciated most was Attia’s encouragement to think for yourself and not follow rules dogmatically (ironic, since his cult following has turned this into a dogma). This led to plenty of interesting questions and self-reflections as I read the book.</p>

<p>A couple of the things I have significantly more appreciation for compared to going into the book: the importance of stability exercise (this clicked for me as a core part of the reason I have different pains), protein (always thought it was important, but may be even more important and the reason I’ve not bulked muscle), and dietary control for avoiding metabolic syndrome (which I appear to have a mild version of already, probably from eating massive meals and a high-carb diet).</p>

<p>It’s also a refreshing perspective on medicine that I’d encourage new residents/med students to read.</p>

<p>Finally: I’m fascinated by the ways in which people can expose society to a new idea in the zeitgeist. Attia managed to popularize the notion of longevity in the mainstream. What a remarkable feat!</p>

<hr />

<p>You can see more about all of the books I’ve read <a href="https://www.goodreads.com/review/list/116173790?sort=rating&amp;view=reviews">here</a>.</p>]]></content><author><name>Sahaj Garg</name></author><category term="blog" /><summary type="html"><![CDATA[Below I share books that have shaped me in different ways: Most influential Most moving Most delightful and fun to read Most useful]]></summary></entry><entry><title type="html">Experimenting with Identity and Personality</title><link href="https://sahajgarg.github.io/blog/experiments/" rel="alternate" type="text/html" title="Experimenting with Identity and Personality" /><published>2024-04-17T21:52:00+00:00</published><updated>2024-04-17T21:52:00+00:00</updated><id>https://sahajgarg.github.io/blog/experiments</id><content type="html" xml:base="https://sahajgarg.github.io/blog/experiments/"><![CDATA[<p>The most important technique I’ve learned in the last year is how to run an identity/personality experiment. It’s unlocked an ability to intentionally learn and grow far greater than any reading, conversation, class, or singular experience. It’s the only way I’ve found to create lasting changes at a core, underlying level to my personality. Many resources I’ve read talk about specific experiments to run, but almost nothing I’ve read talks about how to develop this muscle so you can see it, train it, and use it on your own. This, to me, is the key to unlocking the wisdom well-known in history; it’s how we can speed up the process of learning from our experiences.</p>

<p>So, what exactly is an experiment? Let me give you an example. Earlier last year, I didn’t give my team members concrete feedback if I thought they were doing great work, but could do exceptional work. I wanted to be a “good” manager — kind, supportive, and responsible — and I told myself I’d be irresponsible if I gave someone feedback without 100% conviction in the criticism. I was making two assumptions: (1) if I gave a high-performer feedback, they would be upset and therefore less productive, and (2) if I gave them imperfect feedback, they might adapt to a behavior that I didn’t actually desire. Reading this now, my reaction is “of course that’s not true!” but to actually internalize that, I had to run a series of experiments to test those assumptions. To start, I just gave a specific, concrete piece of feedback to a team member (let’s call them David) on how they might communicate more effectively — sharing my observations of the impact of their current communication style, and how they might be more effective at achieving their goals with some small modifications. My first assumption was challenged nearly immediately: David was extremely grateful and energized to receive the feedback! Because I shared it with compassion and care, not blame, David received the feedback as a gift and as a challenge. With a couple more of these experiments, I mapped out a pattern: high performers want to improve, and receiving feedback (even when they’re doing a good job) is received as an extremely rare gift. Challenging my second assumption took some more time and some more experiments. With David, over the span of several months, he incorporated the feedback in his own way — taking the parts of what I said that were good, and bringing new parts from his own identity and experience — avoiding any harmful consequence. On the other hand, in certain experiments when I shared feedback, I learned why someone was making the tradeoffs and decisions they were — helping me learn how to work more effectively with them without them having to change what they were doing in a counterproductive way.</p>

<p>I call these actions “experiments” for several reasons: I’m testing a hypothesis/assumption, I’m deviating from my default identity and its ensuing default behaviors, I’m often nervous doing them, and I’m creating space for failure or unexpected outcomes. Experiments are bite-sized chunks of things I would not normally do, because of some long-standing assumption or belief that’s holding me back. By making them small enough, I’m able to test whether those assumptions come true without bringing too much risk to myself. At the same time, they’re very intentional and deliberate: I know the outcome I’m worried about, I have an expectation of what might happen instead, and I can check to see what actually happened.</p>

<p>Experiments are effective because they help build wisdom, not knowledge, by challenging assumptions and rewiring your brain. If you’re like me, there are probably lots of things that you “know” to be true, but don’t necessarily believe deep down. I’ve certainly felt this way about many sources of deep wisdom (religious teachings, Buddhism, spirituality, mindfulness) and even self-help resources. They all tell us very similar facts and observations that we might know to be true, but we still have a hard time bringing into our identities and actions. It’s why people often feel a spurt of change as they read a self-help book, acting better for the week they are reading it, and quickly revert back to normal afterwards. The key question is how we internalize that knowledge and turn it into wisdom.</p>

<p>Big assumptions are what get in the way of turning that knowledge into wisdom, because you can cognitively “know” they are not true, but you are likely to <em>take them as true</em> in your actions (and especially non-actions). Based on our many years of lived experience, we’ve told ourselves stories about how the world works - what might happen if we do one thing versus another. Those assumptions, even if we “think” they’re wrong and can reject them on face value, hold us back from internalizing new ways of seeing the world, from turning our knowledge into wisdom.</p>

<p>The standard way people turn knowledge into wisdom is through experience and age; experiments are merely a way to shortcut that. As we live, we experience many data points that tell us what happens when we act in different ways. If we’re observant, we can gradually build a better and better mental model of the world. It’s why we use the word “wise” to generally refer to someone older - like an old, wise grandparent. The goal of experiments is to find a way to build this wisdom <em>actively</em>, instead of passively through experience. Instead of waiting for a data point that challenges our assumptions, we can recognize our agency and take actions that might challenge our assumptions! This creates the opportunity to rapidly rewire your brain.</p>

<p>Experiments can take many shapes and forms: I’ve run experiments to become more assertive, and to find ways to lead from emotion and not just thoughts, among others. For example: I’m naturally inclined to make consensus driven decisions and convince others of my reasoning so I can get buy in. However, sometimes people need to disagree and commit. A deep discomfort would “hold” me whenever I asked  someone to disagree and commit to my perspective — especially if they had many years more experience than me. I worried that I might be wrong and they might resent me for it; I worried that they wouldn’t effectively execute those decisions. When framed like this, the experiment became obvious: I could run a test to disagree and commit once or twice, and see what happened. It was surprisingly effective!</p>

<p>In another experiment, I tested alternatives to my conventional leadership style — logical and analytical. I’ve run a series of experiments where I lead from emotion instead: when I’m experiencing a strong feeling, instead of voicing my opinion, I voice my feeling and its source. For example: “I’m feeling scared that we won’t hit X target because of Y reason.” I had never tried expressing this feeling; my default was to focus on the problem or solution. Expressing those feelings was an experiment with no known outcome: I was uncomfortable doing it purely because I had no idea what would happen! It has turned out to be extremely effective: either I learned that I didn’t need to hold the worry (and was able to avoid sending people off working on an unnecessary solution), or others proposed a solution that was even more effective than what I had in mind, with 100% buy-in.</p>

<p>A tremendous challenge to running experiments is finding opportunities for them. That’s hard because some of our assumptions and behaviors are so ingrained that we don’t even see them, and because experiments often require a situation or context involving other people. I’ve found a number of shortcuts to identify experiments to run “in the moment” — then there’s hardly enough time to even worry about running them! I use one key rule of thumb: <em>never avoid something out of fear or hesitation</em>. There’s plenty of reasons not to do something, but I strive for fear or hesitation to rarely be the reason I avoid an action. Whenever I sense that little nagging thought: “oh no - what will happen if I do X,” I lean in and ask myself, “do I actually want to do it, and what’s the consequence I’m afraid of?” I’ve found many variants of this rule to be effective in helping me recognize potential experiments. I often refer to Claire Hughes Johnston’s operating principle of “say the thing you think you cannot say” as a specific type of fear. The number of times that’s come in handy has been unbelievable! One example is when I’ve noticed a presenter wondering why they’re doing the work they’re doing - calling that out, however uncomfortable it feels - has been extremely effective.</p>

<p>In addition to noticing situations as they arise, it’s also possible to plan for experiments by leaning into your agency. If there’s something I want to happen as a leader that’s not actually happening, I figure out experiments to see how to make that possible (or, of course, recalibrate my expectations). Being high agency and knowing what you want, especially when that’s lingered for some time, is a great way to identify experiments.</p>

<p>Experiments are naturally scary. It’s normal to wonder “what if they don’t work?” or “what if I mess up everything?” The key to avoiding big risks is to make small experiments — for example, try making one or two different comments in a conversation.  If you do that, with genuine intentions to help yourself and those around you, the worst that happens is a need for repair if the experiment goes poorly. Typically a genuine apology with an expression of your intentions will be more than enough. In other cases, people around you might need some time to adjust to your new way of acting. Take it easy, take small steps, and they will accommodate over time.</p>

<p>The biggest risk to experiments is actually with yourself: by making too big of an experiment, by stepping out too far over your skis, it’s easy to fall flat on your face and reinforce the fear you already had. With all experiments and new approaches, there’s effective and ineffective ways to go about them. If you try to make too big of a change too fast, you could easily make it in an ineffective way, or the people around you (especially longstanding, close relationships), could completely reject it. That can reinforce a fear of running experiments themselves, completely stunting growth. Remember to take it easy and start small.</p>

<p>Keeping experiments small has many other benefits: the risk is low, the frequency of experiments can be higher, the growth compounds, and others can slowly adapt to your changes. Opportunities for small experiments show up all the time – basically whenever we slightly hesitate – which for me happens many times a day. If we can take tiny steps that help us change by 1% every day, without taking a large step and being set back, compound growth can lead to tremendous change in just weeks! On top of it, it’s challenging to experiment with big changes around people we already know well. Others have expectations for how we act, and inhabiting another identity altogether can place too much strain on the relationship. Small changes, compounded over time, are much less risky for long-standing relationships.</p>

<p>With that, the most effective lesson that I’ve learned is that experiments are highly effective and not that risky! Initially, I worried with each experiment about the possibility of it going wrong. Now I rarely worry about running experiments. Sure, some of them might not have the intended outcome, but it’s almost always fixable, and I always learn something. This meta-lesson has been transformative for me: it’s what enables me to continue running more and more experiments as I discover more of the things holding me back. Getting comfortable with running experiments has been more valuable than any individual experiment itself; it’s a muscle I’ve been able to repeatedly exercise. Now, I have the opportunity to run multiple experiments every day!</p>

<p>These experiments have produced tremendous value for me. They’ve enabled me to see multiple perspectives simultaneously. They’ve let me build deeper connections with the people I care about most. They’ve helped create an incredible sense of humility as I’ve realized how often I’m wrong about the world.</p>

<p>Mentors are a critical part of this journey — it wouldn’t be possible to do it otherwise. They’ve helped me see the blind spots that I wouldn’t have known to experiment with, given me the tools to actually run experiments, and the feedback to be able to more accurately see whether experiments are going well or not. I’ve found this valuable in the form of regular cadence professional coaching, leadership programs such as Leaders in Tech, and of course, weekly confrontations with my very perceptive wife :). While these aren’t necessary to learn from experiments, they can significantly accelerate the process, and give you wonderful company along the lonely and bumpy ride!</p>

<p>As I’ve come to use this tool more and more, I’ve started to wonder if there are more effective ways to do it. Experiments are slow, stressful, and time consuming. How can we make them faster? I’ve found three shortcuts that help. First is the observation of others. When I hold an assumption about how things work, I can often identify several people I know who challenge that assumption and live a perfectly effective life. (An extreme example: “if I don’t exactly follow some very specific religious custom, I won’t be able to live a good life” – notice there are plenty of counter examples!) I can use those observations as case studies to challenge my own assumptions. Second is through simulation: once I’ve built a good enough mental model, I can predict what might happen with no intervention or several different possible experiments, and predict what to choose. And finally, there’s empathy: I ask myself how I would react if someone ran that experiment on me. Would I be appreciative? Angry? This was most useful when creating spaces for vulnerability at work — I was scared to do so because of fear of judgment, but on the flip side I’ve always appreciated whenever people have made space for that!</p>

<p>Of Observational, Simulation, Thought, and Behavior experiments, the Behavior ones have been the most effective of these techniques - they’ve helped rewire my brain in a way that no amount of reading, simulation, or empathy can substitute. Real, lived experience is what changes how you see the world. Self-help books and other resources tell you specific experiments to run, but I’ve found that learning how to run experiments myself is far more effective. I look forward to the many more years of learning that it can bring!</p>

<p><strong>Acknowledgements</strong></p>

<p>I’d like to thank my coach, David Zeitler, for teaching me how to run the <em>Immunity to Change</em> version of behavioral experiments, question my identity, and change in a persistent way. I couldn’t ask for a better mentor along this journey. I’d also like to thank Tony and Terra, the facilitators of the Leaders in Tech retreat I attended, who’ve helped shape many of my perspectives. Finally, Shruti, my wife, and Tanay, my co-founder, have been incredible resources in giving me feedback through all my experiments and being patient with all my mistakes.</p>]]></content><author><name>Sahaj Garg</name></author><category term="blog" /><category term="personal" /><category term="work" /><summary type="html"><![CDATA[The most important technique I’ve learned in the last year is how to run an identity/personality experiment. It’s unlocked an ability to intentionally learn and grow far greater than any reading, conversation, class, or singular experience. It’s the only way I’ve found to create lasting changes at a core, underlying level to my personality. Many resources I’ve read talk about specific experiments to run, but almost nothing I’ve read talks about how to develop this muscle so you can see it, train it, and use it on your own. This, to me, is the key to unlocking the wisdom well-known in history; it’s how we can speed up the process of learning from our experiences.]]></summary></entry><entry><title type="html">Reflections on 2023</title><link href="https://sahajgarg.github.io/blog/2023-reflections/" rel="alternate" type="text/html" title="Reflections on 2023" /><published>2023-12-27T21:52:00+00:00</published><updated>2023-12-27T21:52:00+00:00</updated><id>https://sahajgarg.github.io/blog/2023-reflections</id><content type="html" xml:base="https://sahajgarg.github.io/blog/2023-reflections/"><![CDATA[<ol>
  <li>Experience and knowledge are very different, and both are critical. Experience must be lived — it cannot be taught. This is the ultimate value of rapid decision making (at a company, and in life): you live more experiences, and grow stronger as a result.</li>
  <li>Fear is almost always a bad reason to avoid doing something — instead, when you notice yourself holding back because of fear, lean into that action. You may be surprised: the outcome of the situation may be far different from your expectations, and the exposure itself may help break down the fear. Fear is different from a decision not to do something because of its risks or consequences.</li>
  <li>Being silent and listening is harder than you’d realize. When talking to people, you often only get to their most controversial - and most interesting - thoughts by waiting for longer than fifteen seconds. Often, the best way to receive feedback is not to ask for feedback, but rather ask a question that lets people answer in as open ended of a way as possible, and infer your feedback from their observations.</li>
  <li>Debates should always start with why. When two people can align on why they are solving the problem they are, it becomes 10 times easier to align on what to do and how to implement it. It’s easy to assume that the what/how of a solution is the only way to achieve a goal. When you get to why, you can understand the real things others care about, and find even better solutions much more easily.</li>
  <li>Negotiation is about finding a win-win, which (like above) is only possible when you focus on why. People get caught up in what they want - by understanding why they want it, you can find another solution that gets everyone what they want. The key often to find a third solution, not the ones both parties came in with.</li>
  <li>The most effective way to deliver constructive feedback: (1) share your observations of the current behavior, (2) ask how they percieve that behavior and what might be leading them to do it, (3) then share your observations of what they could do differently, (4) the suboptimal results that happen as a result of those behaviors, including observations of yourself, (5) the better results that would happen with those changes, and how much higher leverage those would make the person. After understanding the experiences that have shaped that behavior, you will learn so much about the other person, and may be able to help show how their behaviors may not be the best way to achieve the goals they have in mind.</li>
  <li>Ruthlessly prune things that are non-critical or not high leverage. It’s far easier to get more done by getting rid of unimportant work than by putting in more hours. Notice when you’re bored; that implies what you’re doing may be removable, since if it’s truly important, it’d be irritating, not boring. For R&amp;D, pruning work by focusing on the most important parts of the most important problems can eliminate 90% of the work you thought you had to do.</li>
  <li>Clearly articulate what you expect, of yourself and of others. It’s easy to assume other people know what you want. When I’m frustrated, it’s almost always because I haven’t communicated what I want or expect. In those cases, of course something else will be happening.</li>
  <li>Sleeping, exercising, focusing, and reading are all extremely effective ways to leverage the time that you do spend working. I can get more output with fewer hours by focusing on these. Books can give you a lens through which you can think about situations you’re already encountering in life; that is the power of knowledge and words.</li>
  <li>We have a greater ability to influence and shape others’ behaviors and worldviews than we might think on the surface. When people close to you challenge you, it’s not an attack, but rather an opportunity to refine a perspective and help shape both your and their thought processes. We are most effective at shaping worldviews once we have heard others’ objections.</li>
</ol>]]></content><author><name>Sahaj Garg</name></author><category term="blog" /><category term="personal" /><category term="work" /><summary type="html"><![CDATA[Experience and knowledge are very different, and both are critical. Experience must be lived — it cannot be taught. This is the ultimate value of rapid decision making (at a company, and in life): you live more experiences, and grow stronger as a result. Fear is almost always a bad reason to avoid doing something — instead, when you notice yourself holding back because of fear, lean into that action. You may be surprised: the outcome of the situation may be far different from your expectations, and the exposure itself may help break down the fear. Fear is different from a decision not to do something because of its risks or consequences. Being silent and listening is harder than you’d realize. When talking to people, you often only get to their most controversial - and most interesting - thoughts by waiting for longer than fifteen seconds. Often, the best way to receive feedback is not to ask for feedback, but rather ask a question that lets people answer in as open ended of a way as possible, and infer your feedback from their observations. Debates should always start with why. When two people can align on why they are solving the problem they are, it becomes 10 times easier to align on what to do and how to implement it. It’s easy to assume that the what/how of a solution is the only way to achieve a goal. When you get to why, you can understand the real things others care about, and find even better solutions much more easily. Negotiation is about finding a win-win, which (like above) is only possible when you focus on why. People get caught up in what they want - by understanding why they want it, you can find another solution that gets everyone what they want. The key often to find a third solution, not the ones both parties came in with. The most effective way to deliver constructive feedback: (1) share your observations of the current behavior, (2) ask how they percieve that behavior and what might be leading them to do it, (3) then share your observations of what they could do differently, (4) the suboptimal results that happen as a result of those behaviors, including observations of yourself, (5) the better results that would happen with those changes, and how much higher leverage those would make the person. After understanding the experiences that have shaped that behavior, you will learn so much about the other person, and may be able to help show how their behaviors may not be the best way to achieve the goals they have in mind. Ruthlessly prune things that are non-critical or not high leverage. It’s far easier to get more done by getting rid of unimportant work than by putting in more hours. Notice when you’re bored; that implies what you’re doing may be removable, since if it’s truly important, it’d be irritating, not boring. For R&amp;D, pruning work by focusing on the most important parts of the most important problems can eliminate 90% of the work you thought you had to do. Clearly articulate what you expect, of yourself and of others. It’s easy to assume other people know what you want. When I’m frustrated, it’s almost always because I haven’t communicated what I want or expect. In those cases, of course something else will be happening. Sleeping, exercising, focusing, and reading are all extremely effective ways to leverage the time that you do spend working. I can get more output with fewer hours by focusing on these. Books can give you a lens through which you can think about situations you’re already encountering in life; that is the power of knowledge and words. We have a greater ability to influence and shape others’ behaviors and worldviews than we might think on the surface. When people close to you challenge you, it’s not an attack, but rather an opportunity to refine a perspective and help shape both your and their thought processes. We are most effective at shaping worldviews once we have heard others’ objections.]]></summary></entry><entry><title type="html">Designing More Effectively</title><link href="https://sahajgarg.github.io/blog/design/" rel="alternate" type="text/html" title="Designing More Effectively" /><published>2020-11-30T21:52:00+00:00</published><updated>2020-11-30T21:52:00+00:00</updated><id>https://sahajgarg.github.io/blog/design</id><content type="html" xml:base="https://sahajgarg.github.io/blog/design/"><![CDATA[<p>Over the last year at <a href="https://luminous.co/">Luminous Computing</a>, I've worked as a product manager and AI lead defining system-level architectures for our AI chip. My approach to design has evolved substantially over the last year, and developing the below principles has exponentially increased my productivity.</p>

<h3 id="1-design-at-the-system-level">1) Design at the system level.</h3>

<p>It's important to consider how an entire system, or product, solves a user's problem, rather than focusing on features. Systems come with design tradeoffs; by focusing on one feature, there may not be engineering resources or time to focus on another, and by improving one metric as an organization, another might take a hit. Think about system level priorities or bottlenecks when picking which features to ideate on, and vice versa: when identifying a new feature, evaluate how it'll impact the overall system.</p>

<p>Computer architecture is always about system level tradeoffs and finding the right balance. The canonical textbook on <a href="https://www.elsevier.com/books/computer-architecture/hennessy/978-0-12-811905-1">computer architecture</a> explains on its first page that its goal is to "demystify computer architecture through an emphasis on cost-performance-energy trade-offs." For example, at a high level, given a fixed total power budget, by expending more power on memory, a system has less power remaining for computing! If designs focus on optimizing just one feature (for example, just the memory in the above example), they may ignore the impact that those features have on other aspects of the system. Designing at the system level is also important in the presence of bottlenecks. For example, if the goal is to reduce the cost of a system, then effective design would involve breaking down the system costs. If the cost is $900 for the processor and $100 for the memory, then all of the focus should be on improving the cost of the processor by 10x! However, if the cost is $100 for the processor and $100 for the memory, then improving the processor by 10x, while sounding impressive, is not nearly as valuable as improving the processor by 2x and the memory by 2x. Focusing on designing at the system level can help avoid these pitfalls.</p>

<h3 id="2-communicate-with-writing">2) Communicate with writing.</h3>

<p>Conversations are effective for brainstorming, but it's difficult to iterate on ideas orally after a couple of meetings. By writing all ideas down in detail, and in particular stating all assumptions for those ideas, it's possible for other people to identify their disagreements and brainstorm improvements to those ideas. Writing is broadly critical for team collaboration: any information kept in your head is useless to anyone else! It's also helpful to clarify your own thoughts; often, by writing, I've identified holes in my own argument and have understood next steps much more clearly.</p>

<p>An emphasis on written documentation has dramatically changed my personal productivity and that of the Luminous team as a whole. Over the last three months, I have written ~100 pages of design documents, explaining in detail my ideas, assumptions, and implications for engineering strategy. The huge volume of writing comes from the progression in the team's thoughts: every time I wrote out an idea, people could pinpoint exactly where they thought that idea could be improved (e.g. an implementation strategy, or a bad assumption, or an optimization). Moreover, by building a culture that requires written design proposals for ideas to be adopted, other team members have written out arguments that allow me to give more effective feedback. In addition to design documents, written notes on customer conversations are critical for an organization. The written takeaways allow everyone on the team to understand in just five minutes many of the details of a two hour conversation. By writing down all customer conversation notes in detail, other team members have then accurately referenced both facts and opinions from those conversations in subsequent product discussions.</p>

<h3 id="3-prioritize-your-time">3) Prioritize your time.</h3>

<p>Working on design can be abstract, so working on the most important parts of the problem first is critical. If working on one part of the problem changes the problem itself, then all other tasks related to that problem would be rendered useless! A wonderful principle for doing this is prioritizing tasks that maximize the rate of information acquisition, outlined by Jacob Steinhardt in this post about <a href="https://docs.google.com/document/d/1KCSXYmInnBrOnFw5y3kQdNluLTYKt-jF1psyviNAeag/edit#">research as a stochastic decision process</a>. Another guiding principle for prioritizing time is trying to figure out why, what, and then how for a product: don't figure out <em>how</em> to implement something until you know <em>what</em> you're building, and don't commit to <em>what</em> you're building until you know <em>why</em> you're building it!</p>

<p>This is a harder one, and it's much more of an art than a science, but improving (and continuing to improve) this skill led to dramatic gains in my productivity. When focusing on design ideation, my goal would be to put forth an idea (often in a 10 page written document!) and <em>then try to kill the idea as quickly as possible__with the goal of using that killed idea to inspire a better one.</em> The way in which I would focus on killing the idea depended on the problem and was a prioritization decision; sometimes it would be by focusing on system design tradeoffs (e.g. if the cost of photonics in a proposed system was 10x the cost of other important components), other times it would be by evaluating an implementation strategy (e.g. maybe a certain approach is unlikely to succeed due to mechanical integrity issues), and on occasion it would be by running a red-team exercise (e.g. how could someone who's not at Luminous do this even better?). Each of these holepunching questions, notably, leads to an improved idea: e.g. how can we use different implementation strategies to reduce the cost of photonics?; how can we solve mechanical integrity issues with better packaging?; and how do we adapt the problem we're working on to make Luminous a unique winner. Picking the right direction to look for killing and reincarnating subsequent ideas is something that I'm constantly working on improving, and is one of the most difficult parts of my work. Prioritizing correctly, though, can lead to the next idea in two days rather than a week, making a designer thrice as effective.</p>

<p>At a company level, prioritizing time translates to <strong>focusing</strong>. It's hard to make <em>every</em> feature of your product better than all your competitors, but you can focus and make a handful of features truly shine. This fits in with designing at the system level; focus on the aspects of the system that are most important, and make those features the best in the world. At the same time, this also avoids stacked risks: if ten things need to go correctly in order for your system to be valuable, one of them can fail pretty easily, and Murphy's law will kick in. For example, if the cost of a system is determined by five components that each cost $100, and you wish to improve the system by 100x, then you run into issues with stacked risks: if you fail to improve any one of the five components by 100x, then your system can improve by at most 5x. This implies five independent failure points. Instead, by focusing on making a couple of specific aspects of the system shine, a large number of independent failure modes can be avoided.</p>

<p>Focusing will require <strong>deciding quickly with incomplete information.</strong> You'll never have more than 80% of the information you need, and waiting for more than 60-70% means you're too late in making the decision. Using heuristics to make decisions is critical, and knowing the (rare) occasions to reject conventional wisdom and be contrarian is equally important.</p>

<h3 id="4-combine-top-down-and-bottom-up-design">4) Combine top down and bottom up design.</h3>

<p>When designing at the system level, there are two ways you could design a system. On one hand, you could identify a user's problem, and then design a system that targets that very problem. On the other hand, you could look at your tools, and propose different systems that you could build, afterwards examining which ideas are the best at solving people's problems. These two approaches can be divided into a top-down vs. a bottom-up approach to system design. Top-down design, i.e. the canonical <a href="https://www.interaction-design.org/literature/article/5-stages-in-the-design-thinking-process">five step design thinking process</a>, really focuses on making sure the problem is correctly identified. However, it's also important to brainstorm ideas first, since some solutions may help identify new problems that people face, or even create new markets for problems that people didn't realize they had today. It's important to use the best of both of these approaches, since only using one of them may leave you blindsided.</p>

<p>Luminous, as a hardware company, is a case where using both design approaches is especially important. We engage in top-down design when doing customer interviews, understanding AI practitioners' workflows, and understanding what bottlenecks they face in the hardware they use today. This allows us to identify the biggest system-level bottlenecks (for example, more compute power, or more networking?), and design so that we can directly address a customer's needs. On the flipside, many hardware innovations come from a bottom up process. We often ask questions about how much juice we can get out of our matrix multiplier units, or how much networking bandwidth we might be able to supply. These are solutions, and they can inform the types of prospective customers we reach out to and the areas in which we search for problems. For a software company, combining the two design approaches would involve combining empathy on customer pain points with feature suggestions from engineers who might know how to build something cool based on the tools they have available.</p>

<h3 id="5-rely-on-yourself-but-work-with-a-team">5) Rely on yourself, but work with a team.</h3>

<p>Designing requires system level thinking, and you can't hand that off; a single person has to have the entire system in their head to understand the tradeoffs incurred by focusing on different features. Understanding each part of the system, however, requires relying on others for details and working collaboratively to answer all the necessary design questions. You will have to work with a variety of different personalities that each approach design differently; for example, some team members may naturally be top-down design thinkers, and others may be bottom-up. Combining those perspectives is critical to being effective.</p>

<p>Relying on yourself and working closely with your team are two central aspects of Luminous's company culture statement. This has manifested in a variety of different ways. For example, engineers on our team realized the importance of gathering market data, and started reaching out to friends directly and then a broader network via cold emails and second degree connections. Having these conversations as team members with engineering expertise made them especially productive, since the goal was to understand the hardware engineering challenges that algorithm developers were experiencing, a problem we could uniquely empathize with. Instead of waiting for introductions to prospective customers, we relied on ourselves and reached out to friends, and after learning from them, worked with other team members' second degree connections to accelerate the process. On the system design side, AI engineers need to rely heavily on the photonics and digital teams to know what's possible, but can combine information about what is possible themselves to define useful products.</p>

<p><em>Acknowledgements – Thank you to Tanay Kothari and Shruti Ramanathan for invaluable feedback on this post, and many Luminous team members, including Mitchell Nahmias, Dave Baker, Tom Baehr-Jones, and Marcus Gomez for helping shape my perspectives on how to approach design.</em></p>]]></content><author><name>Sahaj Garg</name></author><category term="blog" /><category term="personal" /><category term="work" /><summary type="html"><![CDATA[Over the last year at Luminous Computing, I've worked as a product manager and AI lead defining system-level architectures for our AI chip. My approach to design has evolved substantially over the last year, and developing the below principles has exponentially increased my productivity.]]></summary></entry><entry><title type="html">Principles</title><link href="https://sahajgarg.github.io/blog/principles/" rel="alternate" type="text/html" title="Principles" /><published>2019-12-23T21:52:00+00:00</published><updated>2019-12-23T21:52:00+00:00</updated><id>https://sahajgarg.github.io/blog/principles</id><content type="html" xml:base="https://sahajgarg.github.io/blog/principles/"><![CDATA[<p>I spent some time in December 2019 reflecting on my guiding principles. The result:</p>

<h3 id="life-philosophy">Life Philosophy</h3>
<p>I have a single driving motive: to have a <strong>rich subjective life experience</strong>. To me, this consists of experience and sensitivity of awareness to that experience.</p>
<ul>
  <li><strong>Experience</strong>, broad and rich. I wish for a life full of joy, excitement, and sublime transcendence where the future melts away and the present appears before me in its full splendor and beauty. I expect a wide range of negative emotions and consider them a beautiful part of life. I only wish that they are neither persistent nor permanent.</li>
  <li><strong>Sensitivity of awareness</strong>. For nuanced experience, beyond a small set of emotional labels, I would ideally live in perfect awareness of my mental, emotional, physical, and spiritual states, and exist only in the present. Sensitivity makes every emotion richer, and heightened awareness is necessary for sublime appreciation of the world. In addition to awareness of my experiential self, I ought to sharpen awareness of my narrative self, which consists of the perceptions/stories I hold of my own life.
I prioritize actions that either directly provide a rich experience or improve my sensitivity.</li>
</ul>

<p>In my personal life, I prioritize:</p>
<ol>
  <li>Relationships. I intend to develop meaningful connections, support people and be supported, learn from others, and improve my mental models of how people work.</li>
  <li>Personal development and growth (expanded upon below).</li>
  <li>Physical health, which is a precondition for the types of experiences I want. (Chronic sickness, while an intense experience, limits freedom, is persistent, and sucks.)</li>
  <li>Fun! I want to enjoy life.</li>
  <li>Beauty and aesthetics in art and in the world. I’m just starting to explore this, but I think it’s central to being human. Refining my philosophy on aesthetics is an area for growth.</li>
</ol>

<p>In my career, I prioritize:</p>
<ol>
  <li>Intellectual engagement, both by problems and the people I’m around.</li>
  <li>Personal development and growth (expanded upon below).</li>
  <li>An enjoyable and energizing social environment.</li>
</ol>

<p>I care about personal growth along many axes:</p>
<ul>
  <li>Self-awareness: identification of the self, my strengths/weaknesses, and how I relate to people/the world</li>
  <li>Emotional: deepening my experience, managing it better</li>
  <li>Social: reading people better, understanding their needs, supporting them</li>
  <li>Philosophical: challenge my assumptions/worldviews and improve my mental models</li>
  <li>Intellectual: improve my analytical abilities and expose myself to new ideas</li>
</ul>

<p><strong>Why this philosophy?</strong> The major reason for my prioritization of subjective experience is because a lack of free will seems to imply no other source of meaning. My claims on the lack of free will are based on Pereboom’s induction argument and a biophysical view of nature. If we have no free will, then why should we care about our actions and our lives? For me, the answer lies in the beauty of the subjective experience. Even if I have no free will, I still experience my consciousness, and I would never want to give that up. Other philosophies that focus on ambition and success make no sense if we don’t have the free will to control that success.</p>

<p>There appear to be so many other philosophies to choose from: existential, stoic, nihilist, virtue ethics, etc. Each appears to have good justification! However, I think philosophers are primarily providing explanations for the way they were naturally inclined to live, so there is no a priori reason to support one over the other. (This claim as written is unsubstantiated, but then again, this isn’t a philosophical essay.) For example, Socrates said “the unexamined life is not worth living.” To someone with a personality that cannot help but be skeptical and engage in abstract reasoning, this is obviously true, but for many others, self-examination is painful and unproductive. So, to understand my life principles, I looked to my experiences and developed the theory outlined before.</p>

<p><strong>Optimization.</strong> I naturally try to optimize, and I naturally strive to be successful. At the surface, it seems that my life principles are not aligned with these personality traits. However, if I ask the questions “to what end do I optimize?” and “why do I care about success?,” the answers are primarily to have a fulfilling set of life experiences. I now think optimization and striving for success are less likely to lead to rich, fulfilling experiences, so this is a behavior I ought to guide in a more healthy and productive way.</p>

<p>This leads to my statement that “optimization is the death of my humanity.” Optimization is problematic because it causes worry and draws focus to the future instead of the present. Instead, I hope to be purposeful and intentional with my time, without concern for being optimal. Intentionality is compatible with focus on the present and leads to using time well.</p>

<p><strong>Ethics:</strong> Surprisingly, the fact that I follow a moral code isn’t entirely explained by the subjective experience motive. My ethics, in brief, involve actively helping the people who I care about and doing no harm to others. Actively helping people is also explained by my priorities for rich relationships. Avoiding harm acts less as a motivation for choosing my actions, and more as a rule to avoid certain actions. This ethical system is derived from my experience as opposed to a priori moral arguments.</p>]]></content><author><name>Sahaj Garg</name></author><category term="blog" /><category term="personal" /><summary type="html"><![CDATA[I spent some time in December 2019 reflecting on my guiding principles. The result:]]></summary></entry><entry><title type="html">Demystifying “Matrix Capsules with EM Routing.”</title><link href="https://sahajgarg.github.io/blog/capsules/" rel="alternate" type="text/html" title="Demystifying “Matrix Capsules with EM Routing.”" /><published>2017-11-22T21:52:00+00:00</published><updated>2017-11-22T21:52:00+00:00</updated><id>https://sahajgarg.github.io/blog/capsules</id><content type="html" xml:base="https://sahajgarg.github.io/blog/capsules/"><![CDATA[<p>Recently, Geoffrey Hinton, one of the fathers of deep learning, made waves in the machine learning community by publishing a revolutionary computer vision architecture: <strong>capsule networks</strong>. Hinton has been pushing for using capsule networks since 2012, after he first revolutionized the use of Convolutional Neural Networks (CNNs) for image detection, but only now has he made them feasible. The initial successful approach, published two weeks ago, is titled “Dynamic Routing Between Capsules.” Dynamic routing - which we’ll be exploring in depth throughout this post - allows networks to more intuitively understand part-whole relationships.</p>

<p>In the three days following the release of this paper, another paper on dynamic routing in capsule networks was submitted for review to ICLR 2018. This paper, titled “Matrix Capsules for EM Routing,” is widely speculated to have been authored by Hinton, and discusses a revolutionary new method for dynamic routing - even compared to his first paper. Although this approach shows significant promise, there’s been relatively less media attention surrounding it thus far. Because of the uncertainty of authorship, I refer to “the authors” for the rest of this post.</p>

<p>In this blog post, I’ll be discussing the intuition behind dynamic routing for matrix capsules and the mechanism for dynamic routing. The first part will outline the intuition for capsule networks and how they leverage dynamic routing to be more effective than traditional CNNs. The second part will take a deeper dive into the mathematical and technical details of <strong>EM Routing</strong>.</p>

<p><em>(Note: If you’ve read Sabour et. al., “Dynamic Routing Between Capsules,” the intuition is quite similar but the routing mechanism is very different.)</em></p>

<h2 id="whats-wrong-with-cnns">What’s Wrong with CNNs?</h2>

<p>Since 2012, when Hinton co-authored a paper introducing AlexNet, a CNN that performed better than any previous method on image classification for ImageNet, CNNs have been the standard method for computer vision. In essence, CNNs work by detecting the presence of features in an image, and using the knowledge of these features to predict whether an object exists. However, CNNs only detect the existence of features - so even the image to the right, which contains a misplaced eye, nose, and ear, would be considered a face because it consists of all the necessary features!</p>

<figure class=""><img src="/assets/images/capsules/faces.png" alt="CNN misclassifies bogus faces." /><figcaption>
      A CNN would classify both images as faces because they both contain the elements of faces. <a href="http://sharenoesis.com/wp-content/uploads/2010/05/7ShapeFaceRemoveGuides.jpg">Source</a>.

    </figcaption></figure>

<p>The reason for CNNs failure to deal with such images is the reason why Hinton has criticized CNNs for the last several years. The problem boils down to the way CNNs pass information from layer to layer: <strong>max pooling</strong>. Max pooling examines a grid of pixels, and takes the maximum activation in that region. This procedure essentially detects whether a feature is <em>present</em> in any <em>region</em> of the image, but loses spatial information about the feature.</p>

<figure class=""><img src="/assets/images/capsules/maxpool.png" alt="Max pool." /><figcaption>
      Max Pooling preserves existence (maximum activation) in each region, but loses spatial information. <a href="https://upload.wikimedia.org/wikipedia/commons/e/e9/Max_pooling.png">Source</a>.

    </figcaption></figure>

<p>Capsule networks attempt to address the limitations of CNNs by capturing <em>part-whole relationships</em>. For example, in a face, each of the two eyes is part of a forehead. In order to understand these part-whole relationships, capsule networks need to know more than just whether a feature exists - it needs to know where the feature is, how it is oriented, and other similar information. Capsule networks succeed in ways that CNNs have not because they successfully capture this information and route information from parts to their wholes.</p>

<h2 id="so-what-exactly-is-a-capsulenetwork">So, What Exactly is a Capsule Network?</h2>

<p>A traditional neural network is made up of neurons. Capsule networks impose structure on a CNN by grouping neurons into modules called capsules.</p>

<p><em>Capsules</em> are groups of neurons whose output represents different properties of the <strong>same</strong> feature. By grouping neurons, the output of a capsule is a 4x4 matrix (or 16 dimensional vector), whose entries represent information such as the x and y-coordinates of the feature, the angle of rotation of the object, and other characteristics. For example, in digit recognition, one layer of capsules correspond to digits. One capsule in the output layer may correspond to a 6. Different entries in the 4x4 matrix will correspond to different properties of the 6. As seen below, if you modify one entry of the matrix, the 6 changes in scale and thickness, and if you modify another entry of the matrix, the top hook of the six changes.</p>

<figure class=""><img src="/assets/images/capsules/capsule-features.png" alt="Changing a feature in a capsule." /><figcaption>
      This image, from Hinton’s first paper, shows what happens to numbers when perturbing one dimension of the capsule output. This demonstrates spatial information about the numbers stored in each entry. <a href="https://arxiv.org/abs/1710.09829">Source</a>.

    </figcaption></figure>

<p>When thinking of neural networks, we think in terms of connections between neurons. Although capsules are made up of neurons, it’s easier (and more effective) to think of the network as made up of components called capsules. The entire <strong>capsule network</strong> consists of layers of capsules. Each layer may consist of any number of capsules (often 8–32).</p>

<p>It is intuitively helpful to compare capsules to neurons in a traditional neural network. A single neuron in a NN outputs an activation from (0,1), and represents the existence of a feature. On the other hand, a capsule outputs a logistic unit (0,1) about the existence of a feature <em>as well as</em> a 4x4 matrix representing additional information about the feature. This 4x4 matrix is called the <strong>pose matrix</strong> because it represents spatial information about a feature. The diagram above illustrates what happens when you modify entries in the pose matrix.</p>

<p>In addition, we can consider neuron-neuron connections versus capsule-capsule connections, which would be represented by an single arrow in a typical diagram. In a NN, one neuron is connected to the next using a scalar weight. On the other hand, the connection between two capsules is now a 4x4 <strong>transformation matrix</strong>, because both the inputs and outputs of a capsule are 4x4 matrices!</p>

<p>So far, capsule networks haven’t done anything revolutionary: after all, capsules are often just convolutional layers. The effectiveness of capsule networks comes from the way information is transferred from layer to layer - what’s called a <strong>routing mechanism</strong>. Unlike CNNs, which route information using max pooling, capsule networks use a new process called <strong>dynamic routing</strong> to capture part-whole relationships.</p>

<h2 id="dynamic-routing-through-agreement-from-part-towhole">Dynamic Routing Through Agreement: From Part to Whole</h2>

<p>In this section, we’ll consider how to route outputs from layer <em>(l)</em> to layer <em>(l+1)</em>. This procedure replaces max pooling from CNNs. Remember, as we discussed previously, our network learns transformation matrices (analogous to weights) as the connections between capsules. However, in addition to using these transformation matrices to route information from layer to layer, each connection is also multiplied by a <strong>routing coefficient</strong> that is dynamically computed. Don’t worry if this isn’t clear yet; we’ll first move through the intuition behind these routing coefficients, and then discuss the math later.</p>

<p>To make our analysis more concrete, we consider the example of digit recognition, and only classify the numbers 2 and 3. Our hidden layer <em>(l)</em> consists of a capsule layer with six capsules, corresponding to features, and our final layer <em>(l+1)</em> consists of 2 capsules, corresponding to the numbers 2 or 3. The features in layer <em>(l)</em> represent different components of the numbers 2 and 3. Suppose, for the sake of illustration, that our six capsule nodes to correspond to the following features. A 3 is composed of the first four features, and a 2 is composed of the first, fifth, and sixth features. In our example, our network will have taken an input image of 3, and will try to classify it.</p>

<figure class=""><img src="/assets/images/capsules/routing.png" alt="Dynamic routing." /><figcaption>
      Capsules are shown as nodes, but output matrices. Arrows represent transformation matrices. Red capsules are active, and black capsules are inactive. The green capsules in the output layer have not yet been predicted.

    </figcaption></figure>

<p>Intuitively, we know that every object in layer <em>(l)</em> comes from one of the higher-level objects in layer <em>(l+1)</em>. We know that the first capsule is active in our network because it is either a 2 or a 3. <strong>Our goal is to figure out the probability that our feature is activated because of each feature in the next layer.</strong> Based on why we think capsules in layer <em>(l)</em> are being activated, we can direct our information from layer <em>(l)</em> to <em>(l+1)</em>.</p>

<p>The question becomes how we decide these probabilities. Let’s take the example of the first capsule feature, which could be a part of a 2 or a 3. The number that it comes from depends largely on the spatial relationships between the first feature and other features in the layer. In our example, layer <em>(l)</em> sees the first feature (in the image) above the second feature above the third feature above the fourth. <em>(Note: our pose matrix outputs can capture these spatial relationships between the features).</em> So, layer <em>(l)</em> has seen all the information about a 3 with the correct relationships. If so, then there is significant agreement among layer <em>(l)</em> that a 3 is present. At the same time, the two other nodes for 2 are not active, so there is relatively less agreement that there is a two in the picture. Based on this agreement, we can say with high probability that the first feature actually came from a 3, not a 2.</p>

<p>This overall concept is called <strong>routing by agreement</strong>. By doing this, information from capsules in layer <em>(l)</em> is only supplied to layer <em>(l+1)</em> if the feature appears to have come from that capsule in the subsequent layer. In other words, the 2 capsule <em>only</em> receives information from the first feature if there’s more agreement about the 2. This procedure of dynamically computing which higher-level capsules to send information to is called <strong>dynamic routing</strong>. By dynamically routing the first feature (top hook) only to the 3, we can be more confident about our prediction and assign low probability to a 2.</p>

<figure class=""><img src="/assets/images/capsules/transformation.png" alt="Transformation matrices." /><figcaption>
      Transformation matrices capture the fixed relationships between part and whole.

    </figcaption></figure>

<p>Here’s the trickiest part of dynamic routing: these routing coefficients are not learned parts of a network! The next three paragraphs will explain why. Our learned weights in capsule networks are the transformation matrices - analogous to weights in a NN. These transformation matrices can capture how each feature is related to the whole. For example, this can teach us that the first feature is the top of a 3 or the top of a 2, that the second feature is the middle-right of a 3, and so on.</p>

<p>However, we don’t know whether the first feature is part of a 3 or a 2, given the individual image. If there’s significant agreement that the different parts of a 3 relate properly to the whole, then we think the first feature should be routed to a 3. Other times, though, we might think its part of a two. As a result, it depends on the <em>individual image</em> whether the first feature is part of a 2 or a 3. These routing coefficients are consequently not fixed weights; they are learned <em>every single forward pass</em> of the network, and depend on the individual image. This is why the procedure is called <strong>dynamic</strong> routing - it is part of the forward pass.</p>

<figure class=""><img src="/assets/images/capsules/jumbled.png" alt="Jumbled 3." /><figcaption>
      A jumbled 3. It has all the features of a 3, but is not correctly aligned.

    </figcaption></figure>

<p>Let’s consider how this fixes the example when features aren’t in the right places, which had caused a problem for CNNs. The dynamic routing process is what helps fix this issue. Consider when the four features of a 3 are randomly placed in an input, as to the left. A convolutional network would consider this a 3, because it has the four components of a 3. However, in a capsule network, the “3 capsule” would receive information from the hidden layer saying that the parts of the 3 do not properly relate to form a 3! So, when the first capsule is deciding where to route its output, it does not see agreement for either the 2 or the 3, and sends its output equally. As a result, the capsule network does not predict either a 2 or a 3 with confidence. Hopefully, this shows why the routing must by <strong>dynamic</strong>, because it depends on the <strong>poses</strong> of the features activated by the capsule network.</p>

<p>As we showed above, this routing is only possible because the output of each capsule is a 4x4 matrix. These matrices capture the spatial relationships between the different parts of a 3, and if they do not agree correctly as parts of a whole, the capsule network does not route information with confidence. On the other hand, with CNNs, scalar activations cannot capture the relationship between each part and the whole.</p>

<p>Another way to frame dynamic routing is the concept of <strong>voting</strong>. Each capsule in layer <em>(l)</em> can vote which capsule in layer <em>(l+1)</em> it thinks it came from. The capsule gets a total of one vote, and can distribute this vote fractionally among all subsequent layer capsules. This voting mechanism ensures effective routing of information based on spatial relationships. The voting/routing mechanism comes from the procedure that is discussed below.</p>

<h2 id="em-routing-overallidea">EM Routing: Overall Idea</h2>

<p>We’ve vaguely outlined the intuition behind dynamic routing, but have said nothing about how to compute it. This topic is lengthy and complex, and will be covered in full mathematical detail in the next part of this series. Here, I’ll outline the basic idea behind <strong>EM Routing</strong>.</p>

<p>Essentially, when routing, we are making the assumption that each capsule in layer <em>(l)</em> is activated because it is part of some ‘whole’ in the next layer. In this step, we assume there is some latent variable that explains which ‘whole’ our information came from, and are trying to infer the probability that each matrix output came from the higher-order feature in level <em>(l+1)</em>.</p>

<p>This boils down to an unsupervised learning problem, and we tackle it with an algorithm called <strong>Expectation Maximization</strong>. In essence, expectation maximization attempts to maximize the likelihood that our data (outputs of layer <em>(l)</em>) is explained by the capsules in layer <em>(l+1)</em>. Expectation Maximization is an iterative algorithm that consists of two steps: expectation, and maximization. Typically, for capsule networks, each step is performed three times, and then terminated.</p>

<figure class=""><img src="/assets/images/capsules/convergence.png" alt="EM convergence." /><figcaption>
      The output capsules converge toward the correct classes over the three iterations of EM. <a href="https://openreview.net/pdf?id=HJWLfGWRb">Source</a>.

    </figcaption></figure>

<p>The authors explain the steps are as follows. Expectation assigns probabilities that each capsule in layer <em>(l)</em> was explained by the capsules in layer <em>(l+1)</em>. Then, the maximization step decides whether the capsules in layer <em>(l+1)</em> are active, if they see significant clustering of inputs to it (agreement), and decides what the input is to the capsule by taking a weighted average of the routed inputs.</p>

<p>Don’t worry if this didn’t make sense - I explained a lot of dense math in very few words. If you want to understand the deep technicalities of the routing mechanism, stay tuned for the next blog post later this week.</p>

<h2 id="capsule-network-architecture">Capsule Network Architecture</h2>

<p>Now that we understand the intuition behind capsule networks, we can examine how the authors used matrix capsules in their paper “Matrix Capsules with EM Routing.”</p>

<p>The paper evaluates matrix capsules on the SmallNORB dataset. This dataset consists of five types of children’s toys photographed from different angles and elevations. The dataset accurately captures different <strong>viewpoints</strong> of the same object. Recall that capsules are fundamentally better at understanding viewpoints because of the multidimensional output of capsules (pose matrices)- this is why the SmallNORB dataset is an excellent evaluation metric for the goals of Capsule Networks.</p>

<figure class=""><img src="/assets/images/capsules/smallnorb.png" alt="SmallNORB images." /><figcaption>
      Example images from the SmallNORB dataset. The second row consists of photos taken from a different elevation from the first set. <a href="https://openreview.net/pdf?id=HJWLfGWRb">Source</a>.

    </figcaption></figure>

<p>The architecture consists of a convolutional layer, followed by two capsule layers, and a final classification capsule layer with five capsules, one per class. This architecture performs better than any previous model on the dataset and cuts the previous best error rate by over 45%!</p>

<figure class=""><img src="/assets/images/capsules/arch.png" alt="Network architecture." /><figcaption>
      Architecture in “Matrix Capsules with EM Routing.” See the <a href="https://openreview.net/pdf?id=HJWLfGWRb">paper</a> for more details.

    </figcaption></figure>

<h2 id="advantages-of-capsulenetworks">Advantages of Capsule Networks</h2>

<p>Capsule networks have four major conceptual advantages over CNNs.</p>

<ol>
  <li><strong>Viewpoint Invariance</strong>. As discussed throughout this piece, the use of pose matrices allows capsule networks to recognize objects regardless of the viewpoint from which they are viewed.</li>
  <li><strong>Fewer Parameters</strong>. Because matrix capsules group neurons, the connections between layers require fewer parameters. The model that achieved the best performance only included 300,000 parameters, as compared with the 4,000,000 parameters of their baseline CNN. This means the model can generalize better.</li>
  <li><strong>Better generalization to new viewpoints</strong>. CNNs, when trained to understand rotations, often memorize that an object can be viewed similarly from several different rotations. However, capsule networks also generalize better to new viewpoints because pose matrices can capture these characteristics as mere linear transformations.</li>
  <li><strong>Defense against white-box adversarial attacks</strong>. The Fast Gradient Sign Method (FGSM) is a typical method for attacking CNNs. It evaluates the gradient of each pixel against the loss of the network, and changes each pixel by at most epsilon to maximize the loss. Although this method can drop the accuracy of CNNs dramatically, to below 20%, capsule networks maintain an accuracy of above 70%.</li>
</ol>

<p>In this post, I discuss some details specific to matrix capsules. The details of using matrix capsules with EM routing, instead of using vector capsules in the initial paper, also present three major benefits. I explain the technical details of these three benefits in my next post.</p>

<h2 id="the-future-of-computervision">The Future of Computer Vision</h2>
<p>It’s been two weeks since capsule networks have been introduced, and nobody knows whether they’re another model that people will forget, or if they will revolutionize computer vision. In their current state, they have plenty of drawbacks - most notably the extensive time required to dynamically route using EM from layer to layer. This can be especially problematic when parallelizing over GPUs.</p>

<p>Regardless, this paper effectively acted as a proof of concept for capsule networks, and did so quite elegantly. In the future, though, their efficacy must be tested on a wide variety of datasets, including more applicable ones with many classes, such as ImageNet. Over time, new routing mechanisms may be introduced, and other innovations may make capsule nets more efficient (we’ve already seen this in “Matrix Capsules with EM Routing,” which completely changed the routing mechanism!). If they work, though, the applications are endless - to image classification, object detection, dense captioning, generative networks, and few-shot learning.</p>

<p><em>Acknowledgements</em>  —  A huge thanks to Shubhang Desai, Adi Ganesh, Jeffrey Chang, and Jordan Alexander for giving me excellent feedback on this post. Check out the original papers <a href="https://arxiv.org/abs/1710.09829">here</a> and <a href="https://openreview.net/pdf?id=HJWLfGWRb">here</a>.</p>

<p><em>This post is migrated from the original Medium post, which you can see <a href="https://towardsdatascience.com/demystifying-matrix-capsules-with-em-routing-part-1-overview-2126133a8457">here</a></em>.</p>]]></content><author><name>Sahaj Garg</name></author><category term="blog" /><category term="AI" /><summary type="html"><![CDATA[Recently, Geoffrey Hinton, one of the fathers of deep learning, made waves in the machine learning community by publishing a revolutionary computer vision architecture: capsule networks. Hinton has been pushing for using capsule networks since 2012, after he first revolutionized the use of Convolutional Neural Networks (CNNs) for image detection, but only now has he made them feasible. The initial successful approach, published two weeks ago, is titled “Dynamic Routing Between Capsules.” Dynamic routing - which we’ll be exploring in depth throughout this post - allows networks to more intuitively understand part-whole relationships.]]></summary></entry></feed>