Thought piece
29.9.2026

Building an AI Operating System for Venture Capital

In late January, I wanted to understand what the hell an AI agent actually was, and how it could improve me and the way I work. I started experimenting on Claude Cowork, building a few simple skills and agents around my own workflows.

One became several, and the deeper I went, the more obvious it became that this was not just about automating tasks. If I wanted to get real value from agents, I would have to rethink how I worked in the first place.

The first thing I had to unlearn was my expectation of how software should behave.

INISIGHT 01
Probabilistic systems require a different mental model

Until then, I was used to working with deterministic software — SaaS tools — and with human beings. Working with probabilistic models was very different. The agents wouldn’t consistently produce the results I expected, unlike traditional software, and despite being characterized as intelligent, they could lack common sense, unlike most intelligent people.

A few successful executions therefore meant very little. With probabilistic systems, you need to understand where they are likely to break, which failures actually matter, and build the right checks around those points.

Within the first two months, I had already built and was operating around seven agents. They ranged from automating parts of my day-to-day workflow and helping me keep track of tasks to sourcing strong leads across the main platforms I was using.

As I built them around different parts of my work, another thing became clear.

INSIGHT 02
The best agents become increasingly specific to their user

AI agents need to adapt deeply to the workflows of their users. Two people may be doing the same job but differ significantly in how they gather information, make decisions, use tools, prioritize tasks, and judge whether an output is good enough.

This means that simply taking an agent built by somebody else will rarely produce the best result for you. The underlying logic or architecture may be reusable, but the agent will usually need to be adapted to your own workflows, context, preferences, tools, and decision-making process. The more deeply it adapts to the way you actually work, the more useful and valuable it becomes.

At that point, I started seeing the first meaningful positive signals. The automations were saving me time, but more importantly, they were also improving the quality of my work: better-prepared calls, stronger analysis, and more useful support to founders.

At the same time, another problem was becoming obvious. I was spending far more time than I wanted fixing issues every time an agent broke, behaved unexpectedly, or produced an inconsistent result.

That was when another thing started becoming clear.

INSIGHT 03
Being a strong people manager may sometimes be a handicap

Managing AI agents will become a distinct management discipline, requiring a very different skill set from managing people. In fact, being a strong people manager may sometimes be a handicap, because it creates ingrained expectations about how work is delegated, supervised, and corrected that do not necessarily transfer to managing agents.

Managing agents well requires something different. You need to understand the underlying job deeply enough to know where the system is likely to break, which mistakes materially affect the final output, and which checkpoints deserve human attention.

The goal is not to supervise every step. It is to design the system so that human attention is concentrated where it creates the most value. In that sense, the best agent managers may be the people who understand the work itself most deeply and can translate that understanding into the right controls, evaluations, and escalation points.

So I stopped adding new things for a while and went back to the foundations. I restructured the existing agents and skills, fixed the bugs and gaps in their architecture, and introduced a much more rigorous process for building anything new or updating the existing structures.
A key part of that process became repeated cross-review. I would use both Claude Cowork and ChatGPT to review the same skill or system, then have each critique the other’s recommendations and proposed fixes. The difference was significant.

Using the two systems to challenge each other’s recommendations produced much more robust skills than relying on either one alone. This was one of the first moments when the project started shifting from simply building agents to thinking about the infrastructure, processes, and quality controls required to operate them reliably.

‍

From agents to memory

The value was there and obvious, but still I felt it could be much more. What if I could give these skills, and every future one, access to all the knowledge I had deliberately chosen to absorb — and all the knowledge I would choose to absorb in the future? What if I could build the ideal digital twin of myself? One equipped with the knowledge I considered most important and useful, with a memory far more powerful than mine — able to retain and retrieve what I had learned whenever it became relevant — and with a “clear head” 24/7 rather than for just a few hours per day?

I started building the Knowledge Base from scratch, capturing knowledge from books, blog posts, and other sources, accounting for the fact that some knowledge remains useful for years while some has a much shorter shelf life. After a few sleepless nights, the Knowledge Base was ready and running with a selected set of books and blog posts, feeding relevant knowledge and insights into the different skills. The quality of the work improved, but something was still missing.

As a next step, I wanted to focus on the value that we provide to our portfolio founders and how we can enhance it. A major source of that value is our network. So I translated mine into a network graph. In this network, people and organizations are the nodes, while interactions (a meeting, an email thread, a text exchange, a call) and roles (works at, founded, invested in) are the edges, each timestamped and carrying its source. The network graph answers the question “who is connected to whom, and how?”.

In order to maximize the value of the network graph, though, I needed to build the Spine. The Spine answers the question “what do I know, and what do I think about this node (person, organization, etc)?”. It holds work-relevant context generated through my interactions and the proprietary information that can appropriately be retained and reused: market intelligence, strong founders and operators, my own judgments and scores, action items, and other relevant signals.

While the Knowledge Base captured what I had chosen to learn from the outside world and considered important, the Spine captured what I was learning and generating through my own work, decisions, and network.

The Spine ended up being the missing piece of this AI puzzle.

‍

From memory to a system‍

At that point, what had started as a collection of individual agents had become an actual system. The agents were still the visible layer, but most of the value was increasingly coming from what sat underneath them: shared memory, accumulated knowledge, proprietary context, and the infrastructure connecting everything together.

At a high level, the architecture looked like this:

The agents are the visible layer. The leverage increasingly sits in the shared memory, knowledge, network, and measurement layers underneath them.

At the top are the agents, organized around the core parts of our work: sourcing, selection and decision-making, portfolio support, and capacity building across the investment team. Underneath them sit the shared intelligence layers (the Spine, Knowledge Base, and Network Graph), which allow information captured in one part of the system to become useful elsewhere. The network alone now represents thousands of professional relationships, together with the relevant interactions, judgments, and signals associated with them.

Until then, each agent had access to the capabilities of the underlying models and, where relevant, the knowledge from the Knowledge Base. The Spine introduced a fundamental change by providing a shared memory and intelligence layer across the entire system.

Underneath that relatively simple idea sat a much harder problem: resolving identities across different sources, preserving where every piece of information came from, deciding what could be reused and by whom, and making sure outdated information did not live forever.

‍

When the system started compounding

From that point on, an insight captured by one part of the system could become useful to another. Knowledge about an exceptional operator could surface when one of our portfolio companies needed exactly that profile. A market signal picked up through my network could inform both a future investment decision and the support provided to an existing portfolio company.

A core part of our job is receiving, processing, and acting on high-quality information at the right time. Unfortunately, our brains, and specifically our memory (especially when we don’t sleep enough), are not capable of fully utilizing every piece of information every time it becomes relevant.

A system that can retain that information, connect it to the right entities, and surface it whenever it becomes relevant can create a significant moat for its user. If that system is continuously enriched by every new interaction, insight, decision, and outcome, and can learn from them, it creates a flywheel in which each new input improves the system and increases the value of what is already there.

Over time, that compounding effect can create a massive delta in judgment, speed, and quality of execution.

Of course, not every piece of information should flow freely through this loop. Personal, sensitive, or confidential information needs to remain appropriately restricted, with access, retention, and reuse governed by its source and purpose. But whenever information could lawfully and appropriately be reused, the system no longer had to depend on me remembering that it existed, where it came from, or when it might become relevant.

In any case, the system is designed to support decisions, not make them autonomously; human judgment remains at the critical decision points.

‍

Now the million-dollar question: Is it worth the pain and time?

A response off the top of my head would be based on anecdotal evidence, like 50% more capacity for calls and meetings, or more connections and value to portcos, but I prefer facts and numbers whenever I can get them.

So I built a measurement layer around the system.

Lineage records every decision, action, and recommendation made by the agents, with timestamps, and tracks what ultimately happens with each of them. This allows me to measure whether the machine is actually generating value, where that value comes from, and which parts of the system are contributing to it.

Using the same logic, I built the Decision Ledger around myself. It records my own decisions, actions, recommendations, and their eventual outcomes over time, allowing me to measure the quality of my work and judgment, and improve both.

Lineage measures whether the machine is worth running; the Decision Ledger measures whether the person running it is right.

‍

What comes next

There is a lot underneath this architecture that I have deliberately skipped here. How do you build memory that knows what to retain, what to forget, and how long information should remain relevant? How do you turn thousands of relationships and interactions into useful network intelligence? How do you evaluate probabilistic agents and know when they can be trusted? And, ultimately, how do you measure whether any of this is actually making you a better investor?

I will dig into some of these individually in the next posts, by which point I should also have more data from Lineage. This one was about the journey from one agent to the system — and the realization that the real value did not come from any single agent, but from connecting them all into something that could learn and compound.

‍

Author: George Karabelas