# Stop using AI. Start having AI systems running.
**作者**: Will Chen
**日期**: 2026-07-23T07:23:07.000Z
**来源**: [https://x.com/stablechen/status/2080191964247625824](https://x.com/stablechen/status/2080191964247625824)
---

I used to use AI the normal way. I’d open ChatGPT or Claude, ask for something, get an answer, and close the window. The next conversation started from scratch.
That’s no longer how I work. I have AI systems running.
When I open my coding agent, it already shows my goals, focus, and habits. One sentence about playing pickleball can update cardio and my social log. I can ask questions against fourteen years of journals, old AI conversations, work recordings, and transcripts from walks. At night I can send agents into separate Git branches and review what they built the next morning.
It looks very planned when I list it all out. It wasn’t.
Most of it came from complaining to a coding agent. I’d lose a good thought on a walk, forget to open a habit app, or find myself correcting the same agent behavior again. I’d say what was annoying me, let the agent change the repo, and keep using the new version. This happened almost every day for a few months.
Each change felt trivial. Together they became an environment that remembers what I’m doing, carries state from one day to the next, and keeps working after I leave.

I have one fairly stubborn rule: if using a system feels like extra work, I’m eventually going to stop using it. I don’t find that morally interesting. I’d rather change the interface, timing, or division of labor than keep rediscovering that I can’t force myself through the same bad setup forever.
AI makes it cheap to keep changing the setup. Language practice can happen during cardio. Thinking can happen while I walk. A habit tracker can live inside the coding agent I already have open. A correction I make once can become a rule that every later agent inherits.
This is the stack I ended up keeping.
## Pi: the harness that grows the stack
I have a hard time getting excited about most AI productivity apps. There are so many of them, and I’m usually thinking that somebody else’s hand-built workflow isn’t going to fit me as well as my own repo plus Claude. The tools I keep tend to add a new primitive or extend what a coding agent can already do.
Pi is the main harness for AI Organs. I didn’t feel the x-factor with OpenClaw, even when everybody was excited about it. I felt it with Pi, t
The reason is simple: Pi can modify the environment it’s running inside. If I want another command, a different view, or a rule that should survive the conversation, I tell the agent in normal language and it edits the harness. I don’t file myself a feature request and come back on the weekend.
This still feels like firing off Claude. The difference is that the answer sometimes becomes part of the place where I work. I asked for my habits to appear instantly, so Ctrl+H now opens them. I wanted one log entry to update several parts of my life, so the logging command learned how. If I have to give the same correction twice, I tell the agent to put it in an instruction, test, or example.
Everything lives in one repo. AGENTS.md is symlinked to CLAUDE.md, so I can drive the same system with Claude Code when I want Claude Fable 5, or through AI Organs when I want Pi. I care more about the repo than model continuity. The files, commands, history, tests, and permissions are still there when I switch models.
This is why the stack grew without feeling like a separate project. A problem appears while I’m already working. I complain in the same window, the agent changes the place where it happened, and I go back to what I was doing.
After months of that, the repo contains a strange amount of me: commands I reach for without thinking, policies I no longer want to reconsider, interfaces fitted to my attention, and records that let agents begin with context I’d never fit into a prompt. None of this felt ambitious while I was adding it one irritation at a time.
I started calling the stateful pieces organs because they don’t behave like ordinary apps. Each one owns a domain—habits, money, language learning, work history—and leaves behind enough state for an agent to pick up later. I can switch models without losing that accumulated body.

## Tools
Some of the most important pieces are ordinary products used in a slightly weird way.
## Plaud NotePin
The Plaud NotePin is a $180 recorder I wear on a lanyard when I walk through Golden Gate Park or along Ocean Beach. The price is honestly offensive because the hardware is basically a Bluetooth microphone with a button, and then they charge monthly for transcription. What the fuck are you paying $180 for?
I bought it anyway. A phone app can do the same thing for most people, but I wanted to record for hours and send the output into my own pipeline. The dedicated object removes the choice. I put it on, hold the button, feel it buzz, and start walking.
On a long day I can walk 40–50,000 steps because I’m busy talking. I don’t need a subject before I leave. I can spend the first ten minutes complaining that I have nothing to say, repeat myself, take the wrong branch, and eventually find the thing I actually care about. I keep walking because I haven’t finished the thought.
The recording quality doesn’t have to be great. Wind, traffic, and half-finished sentences used to make this kind of archive pointless. Modern transcription is spectacular. It can recover words that are barely audible and an LLM can reconstruct an argument that arrived out of order.
People also wildly overestimate how much text a human life produces. Hours of speech sound like an absurd data problem, but it isn’t that many tokens. It’s easily managed.
This is why I’m willing to ramble. Coherence can happen later. A model can recover the argument, keep the phrases that still sound like me, show me what I kept circling, and ask whether it got the reconstruction right. The only part I have to do live is capture the thought before it disappears.
Andrej Karpathy recently described the ten-minute desk version of this pattern: switch to voice, ramble without filtering, and let the model reconstruct the thought. He writes that “LLMs are somehow very good at reconstructing long incoherent rambles.”
https://x.com/karpathy/status/2079610838143623371?s=20
Plaud is the same move worn on a lanyard and allowed to run for an afternoon.

Since April 14, the NotePin has captured 61 transcripts and 586,192 words across 41 days. Those walks now feed essays, language lessons, decisions, prototypes, and searches I couldn’t have anticipated while I was talking.
You’d be surprised how much signal you emit in daily life. Most of it disappears because it doesn’t arrive in publishable order. Walking and commuting used to be dead time. Now they produce raw material while also making the day more enjoyable.
The pipeline still lags sometimes. I don’t care. A transcript can be summarized next week or reprocessed by a better model next year. The only permanent failure is not recording it.
## OBS + timer
One thing I’ve always struggled with is starting remote work. You can call it procrastination or ADHD, but I don’t find either label very helpful. They make it sound like a moral failure and don’t tell me what to change. The actual problem is usually the ten minutes before I begin. Once I’m working, I’m fine.
I open OBS, the free streaming app, and record my screen with a small box of my talking head in the corner. Then I set a physical timer, usually for 45 minutes. That is the backbone of my ability to work.
I’m basically pretending to livestream. I explain what I’m about to do, narrate while I do it, and talk through whatever is confusing me. The audience is my future self, plus whatever agents end up reading the transcript. I like to ramble, so this makes the work much more engaging than sitting silently at a computer trying to be disciplined.
I also make weird noises and occasionally sing to entertain or embarrass my future self. Once you get over how strange it is to have a tiny livestream with no viewers, I can’t imagine working without it.

The recording fixes another problem I didn’t have a good name for: work often feels like it won’t count. You can spend an hour trying something, get nowhere, close the laptop, and feel as if the hour disappeared. If I record it, the session exists. Even a bad block shows what I tried, where I got stuck, and what I was thinking at the time.
The files upload to Dropbox and a private, unlisted YouTube playlist. YouTube is free video storage. I have gigabit upload at home, storage is cheap, and my protocol is record, trim the ends, post. If I open Premiere, DaVinci, or CapCut, I’ve already fucked it up. Editing turns a work record into a content project.
The timer matters, but I use it differently depending on the work. For execution, I keep the countdown visible. For open-ended thinking, I hide it; watching the minutes drain makes generative work worse. Either way, the timer externalizes the boundary so I don’t have to keep asking myself whether I’ve worked long enough.
This is seriously a life hack for ADHDers. It’s useful for self-study, auditing where your time went, and holding onto a thread long enough to finish it. Speaking out loud gives the attention somewhere to go.
OBS is reality contact.
I got the idea from watching George Hotz do eight-hour livestreams and stay completely locked in. There are now more than 1,500 hours on my private playlist. A friend once doubted that I was actually working, so I sent him the link. He looked at it and invested $10,000 immediately. Fifteen hundred hours of footage is hard to bullshit.
I used to justify recording everything by saying that a future model might be able to use it. That future has already started. Cortex contains transcripts from 542 sessions, and agents retrieve what I said while I was working alongside my chats, walks, journals, and logs. I’m not saving video in the abstract hope that it becomes useful one day. Five hundred forty-two sessions are already memory.
## Raycast snippets
Raycast is basically a command line for normal people on a Mac. The first thing I tell people is to disable Spotlight and give Command+Space to Raycast instead.
The real difference is programmability. Spotlight searches for things Apple decided it should search for. Raycast can run extensions, scripts, snippets, window commands, and my own AI commands from the same box. I have an extension where I type money and see my bank balances. If I want a new behavior, I can write it once and summon it from anywhere on the computer. It feels closer to adding a command to myself than installing another app.
Snippets are the smallest part and maybe my favorite. I type a short trigger in any text box and Raycast expands it into a timestamp, a template, or a whole way of interacting with an LLM.
My most-used one is `!bear`, the clarity bear:
> First think deeply, reviewing and then interview me, one question at a time presenting me well-considered options and suggestions to advance your clarity until you reach 100% clarity. For each question try to indicate what the source of confusion might be, or how answering this will determine something you need (be concise). Try to opt for multiple-choice when it’s clear enough to do so.
I use it when I don’t know what to prompt because I don’t know what I think yet. The model starts interviewing me one question at a time. Multiple-choice is the trick. It gets the options right about 95% of the time, and picking or rejecting an option is much easier than inventing a clean prompt from a confused state. I use it for stuck projects, specs, braindumps, mental-state debugging, and other people’s dilemmas.
`!time` is even simpler. It expands into a timestamped Markdown line. I pair that with Raycast Notes and hold-to-dictate, which is faster and more reliable than building some elaborate Notion integration. Local and instant beats integrated and networked.
The clarity bear was once a product. I built a whole application around braindump → interview → clarity meter. The product died. The useful part survived as sixty words behind a keystroke. That exact passage has appeared in 69 files since, carrying the same fossilized typo every time. I never retype it.

## Morning mantra + goals.md
The morning mantra is a text file I read out loud before work. I’m on version eight. Calling it a mantra makes it sound spiritual, but I treat it like software: lines get added, reviewed, and deleted when they stop being useful. I recently removed a few because they were just anxiety pretending to be wisdom.
It exists because I can understand something perfectly at night and wake up the next morning without access to the feeling. If I know an insight is going to decay, I put the instruction somewhere tomorrow’s version of me will actually see it. The final line tells me to start the timer.
goals.md stays visible in the agent interface. It contains the two outcomes I’m currently trying to reach and the work that matters now. Agents can explore how to get there, but I don’t want the destination rewritten by whichever conversation happened most recently.
## GPT Live inside cardio
Fuck it: GPT Live is the best language-learning app I’ve used. After one good session my reaction was, why am I not scheduling thirty minutes of this every day while doing cardio?
Language apps optimize exercises. I want to know whether I can explain metacognition in French, argue about a business idea, role-play a weird situation, or flirt with a virtual waitress. This is the first tool that makes me want to talk about anything instead of completing the next lesson.
The obvious alternative is talking to people, but that requires scheduling, confidence, and enough energy to be publicly bad at a language. GPT Live has infinite patience. I can arrive tired, mix languages, change subjects halfway through, and ramble until the sentence appears. It still understands what I’m reaching for.
The important feature is interruption. The instant I don’t know a word, or can’t remember whether something is masculine or feminine, I ask and get the answer inside the conversation. The correction arrives at the exact point of uncertainty. A flashcard app can’t do that, and a class usually won’t.
I run the session during cardio so it has a fixed place in the day. The physical activity is already happening; the conversation makes it interesting enough that I stop watching the clock.
The sessions aren’t confined to one language. They run about 98% in French, with jumps into Cantonese or Korean. If I want to practice a Korean sentence, the model explains it in French, gives me the complete Korean, breaks it into fragments, and waits while I repeat each one. French is becoming the medium of cognition rather than the object being studied.
We rehearse conversations I’ll actually need to have. I can show up halfway through cardio, tired and mixing three languages, and work through something difficult I want to explain to a real person. GPT Live keeps track of the meaning while the languages change.
I was already doing a cruder version in 2024, reading native articles “sentence by sentence, word by word” with voice mode. It became consistent when I attached it to cardio. The sessions now read from and feed the Korean organ’s _state.md, so live practice and structured lessons share a record of what I know and where retrieval fails.

The model says the Korean sentence. Then the first fragment. I repeat it. It adds the next fragment. We keep going until I can say the whole thing without help.
## Organs: letting the state accumulate
The tools above still need me to press record or start the conversation. The organs are code and files that stay available inside the agent. They keep enough history that I don’t have to explain the domain again every time.
## `organs log`
I’ve kept journals for years, but journaling is still an activity you have to decide to do. Most ordinary events aren’t important enough to make me open a separate app and write a reflective entry.
I type `organs log` followed by one sentence in the terminal where I’m already working. The command adds a timestamp and appends the sentence to today’s Markdown file. One file per day, no form, no categories to remember. About 112 of the last 121 days are recorded.
If I write that I played pickleball and joined a pickup game, the agent checks cardio and records a social event. I don’t open the habit tracker, then the social tracker, then a journal. I describe what happened once.

The original sentence stays verbatim. Agents love turning specific thoughts into bland summaries, so the rule is to add interpretation without replacing the source. A weird phrase, repeated complaint, or exact order of thought might be the useful part later. Storage is cheap. Re-creating the sentence isn’t possible.
After a few months, I can ask what I was doing the last time I felt this way or whether a new idea is actually new. The answer comes back with dates and my own words.
## Habits
I’ve tried normal habit apps. The problem is that the habit tracker itself becomes another habit. If I have to remember to open a separate dashboard before it can remind me what to do, it has already failed.
My tracker is a SQLite database rendered directly inside the agent interface. Ctrl+H opens it underneath my goals and current focus. If I don’t know what to do next, I press the shortcut and look at what’s still open.
The agent sees the same list. If I log a pickleball game, it can check cardio automatically. If several habits are still open, it can suggest one from the interface I’m already using. The more I use the coding agent, the more often I see my habits. That’s why this tracker works when the others didn’t.

I’m not trying to perform a perfect streak for a screenshot. I want today to feel connected to a larger system I chose. If a habit becomes overwhelming, I can scale it down or skip it without losing the record.
## Chat archive and Cortex
AI conversations are weirdly disposable by default. I can spend three hours developing an idea with Claude and then open a new chat next week with none of that work available.
Now every conversation gets exported to Markdown and summarized. I have 3,520 summaries. In one run, agents processed 3,273 conversations in about two hours, pulled the 77 with the strongest writing material, and found 101 possible drafts.
A chat window is an interface. It shouldn’t determine how long the thought lives.
Cortex indexes 9,480 documents and 189,703 chunks in a 1.77 GB SQLite file: daily logs, fourteen years of journals, conversations, meetings, OBS sessions, and walks. It runs keyword and vector retrieval together, then reranks the results. grep is not very good when the corpus is your entire life and you don’t remember the words you used.
That was the original problem. I’d remember having a thought but not when, where, or how I phrased it. The first time I asked a question against years of my own writing and got an answer with dated evidence, it was honestly a little shocking. A vague intuition suddenly had a history.
Before a decision, I can ask for the last five times I made a similar one. For language practice, an agent can inspect what I actually talk about and generate vocabulary from my obsessions. I don’t have to sit there summoning an autobiographical word list from memory. The curriculum is already present in the record.
But search by itself isn’t that interesting. Everybody has search. What became useful was being able to run operations on my past.
Agents can move chronologically through every log, walk, chat, and journal and extract the ideas into a schema. Another fleet can build prototypes from those objects. Test agents can exercise them. My feedback can change the evaluation criteria, and a tournament can rank what survives.

I can also ask several agents to simulate possible futures using evidence about what I sustain, abandon, enjoy, and return to. It isn’t prophecy. It’s a better thought experiment because the agents can read more of my history than I can hold in working memory.
This is why I think “memory” is too small a name for it. The records can be filtered, compared, turned into schemas, split across agents, used to build prototypes, and scored against my later feedback. The computation I already spent thinking isn’t trapped in the moment when I did it.
I once called Cortex “not a second brain, but a thousand brains.” A second brain helps you find what you stored. Cortex lets many minds work over what you lived.
I trust the archive for evidence, not intent. It can tell me what I said and how often I returned to it. It doesn’t get to decide what I want now.
Anything new I build can start with that history instead of an empty prompt.
## A few other organs
The worklog keeps one Markdown file per day explaining why the code changed. Git preserves the diff; it doesn’t preserve what I was trying to do.
The email organ ran a written triage policy over 33,166 messages and produced 82 action items. The policy also contains a hard rule that nothing I sent can be deleted.
The finance organ talks directly to the Mercury API—@mercury Personal has a great API, shoutout—and traces transactions back to the events that produced them. I don’t use the MCP because it’s a wrapper around an API I can already call.
Tasks stay in Things 3. It has a good phone interface, its schema has barely changed in ten years, and things-cli can read the SQLite database. I’ll replace it when I have a phone interface worth replacing it with.
## Taskbox, Git, and Linear
One Claude Code session barely touches a Max subscription. If I drive it by hand, I’m leaving most of what I paid for unused. Taskbox is how I fan work out into separate agent runs without losing track of which prompt produced which result. It has handled 2,271 runs and more than 300 million tokens.
Git is almost absurdly well suited to agent work. Each agent gets an isolated branch or worktree, changes real files, and returns a diff. If three agents try three approaches, I can compare the code instead of comparing three confident paragraphs in chat.
The pull request is where the branches have to become one accepted version. Tests run, review agents inspect the diff, and I merge what survives. This is basically what Unix was made for.

Linear holds the unresolved question while the agents work. I have a strong opinion about this: Linear is for converging states, not checking shit off a vague task list. A good issue is the home where competing branches, evidence, and review eventually settle into one answer.
Once I select an approach, the reason can become a test, example, or instruction. That’s how one moment of taste changes hundreds of later runs.
/will gives Claude a model of my 58 skills and seven protocols I wrote for myself. Claude Code already knows how to call a shell or read a file; this command tells it what Will Chen can do.
I am using Claude while Claude is pretending to use Will Chen, as an agent.
It identifies the part that still needs human judgment, taste, conversation, or courage, then hands me the context and a procedure I wrote earlier. The agent can often retrieve my own capabilities more reliably than I can.
## Systems: when the day starts to feel different
At this point the stack stops feeling like a collection of tools. I can close the laptop without losing the state of the day. The files remain, the agents keep their assigned work, and the next session doesn’t begin with me trying to remember everything.
## The morning briefing
The morning briefing is the first thing I read. It pulls together my goals, current focus, open habits, tasks, calendar, body data, finances, and whatever the agents finished overnight.
I don’t open six apps and reconcile them in my head. The coding agent and I start with the same screen. If I ask what matters today, it can answer from the actual state instead of generating generic advice.
It isn’t a report I save for later. It tells me what to do next.
## One day through the stack
In the morning I read the briefing and mantra, then start the timer.
During the day, events go into one-line logs. A walk produces a transcript. Ctrl+H shows the next habit. An OBS recording becomes part of the archive. I use each thing normally and the records accumulate behind it.
In the evening I review what happened, update the worklog, and send out work that doesn’t need me. Agents process transcripts, try implementations, run tests, or prepare options. By the time I return the next morning, there is usually something concrete waiting for review.
## The loops
For me, an agent is a loop. It reads a state, does something, checks the result, and either continues or stops. Treating every subagent like a separate employee with a personality only makes the coordination harder to reason about.
A walk becomes a transcript. Cortex retrieves related material. Taskbox sends an operation across it. The result comes back as a file, branch, log entry, or update to another organ. If an agent gets something wrong, the correction becomes a test or instruction before the next run.

Every useful loop touches something outside the model. A completed workout updates the tracker. A bank transaction updates the financial state. A test says whether the code works. Without those answers, the agents would only become better at agreeing with themselves.
The model reads the files, calls the commands, and moves information between domains. I can replace the model because the repo, records, policies, tests, and permissions survive the session.

## From using AI to having AI systems running
I didn’t decide to build a second brain or a life operating system. I kept using a coding agent during my real day and changing whatever annoyed me. Because those changes lived in one repo, each new fix started from the last one.
This is the part I think people miss about personal AI. The architecture can emerge later. The first requirement is somewhere for a correction to remain. Put the tracker where you already work. Record the session if recording makes it easier to start. Move language practice into cardio if that’s where it becomes enjoyable. When the fix works, keep it.
Most AI demos make generation look like the breakthrough. We already generate too much stuff. The larger change for me is that connection became cheap: a walk can feed an essay, a log entry can update a habit, an old conversation can become a prototype, and one correction can reach every agent that comes after it.
In 2024, I wrote: “record as much data as possible—I believe compute will be cheaper in the future and allow us to go over all the notes that we took.” At the time, that was a bet. Now the record can be searched, transformed, distributed across agents, tested, and returned as action.
This essay came out of the same stack. I dictated parts while walking, retrieved the relevant logs and conversations, sent sections through separate agents, checked the claims against dated evidence, and kept revising the result in the same repo.
AI Organs is only a few months old and it has already stabilized a surprising amount of my life and mental energy. An insight has somewhere to go. It can become a command, a policy, a test, a lesson, a draft, or a task for another agent instead of remaining one more thing I’m supposed to remember.
That is the shift. I still use AI, obviously. But more and more often, I have AI systems running.
## 相关链接
- [Will Chen](https://x.com/stablechen)
- [@stablechen](https://x.com/stablechen)
- [5.7K](https://x.com/stablechen/status/2080191964247625824/analytics)
- [https://x.com/karpathy/status/2079610838143623371?s=20](https://x.com/karpathy/status/2079610838143623371?s=20)
- [@mercury](https://x.com/@mercury)
- [Upgrade to Premium](https://x.com/i/premium_sign_up)
- [3:23 PM · Jul 23, 2026](https://x.com/stablechen/status/2080191964247625824)
- [5,776 Views](https://x.com/stablechen/status/2080191964247625824/analytics)
---
*导出时间: 2026/7/23 21:29:24*
---
## 中文翻译
# 别再用 AI 了。开始让 AI 系统运行起来。
**作者**: Will Chen
**日期**: 2026-07-23T07:23:07.000Z
**来源**: [https://x.com/stablechen/status/2080191964247625824](https://x.com/stablechen/status/2080191964247625824)
---

我以前用 AI 的方式很正常。我会打开 ChatGPT 或 Claude,提个问题,得到答案,然后关掉窗口。下一次对话又得从头开始。
这不再是我的工作方式了。我有 AI 系统在运行。
当我打开我的编程代理时,它已经显示了我们的目标、关注点和习惯。一句关于打匹克球的话就能更新我的有氧运动记录和社交日志。我可以针对十四年的日记、旧的 AI 对话、工作录音和散步时的逐字稿进行提问。晚上,我可以把代理派到不同的 Git 分支上,第二天早上再审查它们构建了什么。
当我把这些全部列出来时,看起来规划得很好。其实并不是。
大部分内容都源于我对编程代理的抱怨。我在散步时忘记了一个好想法,忘了打开习惯应用,或者发现自己又在纠正同样的代理行为。我会说出哪里让我烦,让代理修改仓库,然后继续使用新版本。这种情况在几个月里几乎每天都会发生。
每一次改变感觉都很微不足道。但它们合在一起,形成了一个能记住我在做什么、将状态从一天传递到下一天、并且在我离开后继续工作的环境。

我有一个相当固执的规则:如果使用一个系统感觉像是额外的工作,我最终会停止使用它。我不觉得这在道德上有什么趣味。我宁愿改变界面、时机或分工,也不愿一次次地发现自己无法永远强迫自己去忍受同一个糟糕的设置。
AI 让不断改变设置变得很廉价。语言练习可以在有氧运动时进行。思考可以在散步时进行。习惯追踪器可以住在我已经打开的编程代理里。我做过的一次纠正可以成为每一个后续代理继承的规则。
这是我最终保留下来的技术栈。
## Pi:成长的技术栈的挽具
大多数 AI 生产力应用都很难让我兴奋。这类应用太多了,而且我通常认为别人手工构建的工作流不会像“我自己的仓库 + Claude”那样适合我。我保留的工具往往会添加一个新的原语,或者扩展编程代理已经能做的事情。
Pi 是 AI Organs(AI 器官)的主要挽具。即使当大家对 OpenClaw 感到兴奋时,我也没感觉到它的特别之处。但在 Pi 身上,我感受到了。
原因很简单:Pi 可以修改它运行所在的环境。如果我想要另一个命令、不同的视图,或者一个应该在对话结束后保留的规则,我会用正常的语言告诉代理,它会编辑挽具。我不用给自己提一个功能需求,然后等到周末再回来看。
这感觉还是像给 Claude 发指令。不同的是,答案有时会成为我工作场所的一部分。我要求我的习惯能立即显示,所以 Ctrl+H 现在会打开它们。我想要一条日志条目能更新我生活的几个部分,所以记录命令学会了如何做到这一点。如果我必须给出同样的两次纠正,我会告诉代理把它放入指令、测试或示例中。
所有东西都在一个仓库里。AGENTS.md 软链接到了 CLAUDE.md,所以当我想要 Claude Fable 5 时,可以用 Claude Code 驱动同一个系统;当我想要 Pi 时,可以通过 AI Organs 驱动。相比模型的连续性,我更在乎仓库。当我切换模型时,文件、命令、历史、测试和权限仍然在那里。
这就是为什么技术栈的增长感觉不像是一个独立的项目。问题出现在我正在工作的时候。我在同一个窗口抱怨,代理改变它发生的地方,然后我回到我正在做的事情上。
几个月后,仓库包含了奇怪大量的“我”:我不假思索就能触及的命令、我不想再重新审视的策略、适合我注意力的界面,以及让代理能以我无法塞进提示词的上下文开始的记录。当我一次只解决一个烦恼时,没有任何一件事感觉雄心勃勃。
我开始把这些有状态的片段称为“器官”,因为它们的行为不像普通的应用。每一个“器官”拥有一个领域——习惯、金钱、语言学习、工作历史——并留下足够多的状态供代理以后拾取。我可以切换模型而不失去那个累积的主体。

## 工具
一些最重要的部分其实是以一种稍微奇怪的方式使用的普通产品。
## Plaud NotePin
Plaud NotePin 是一个售价 180 美元的录音机,当我穿过金门公园或沿着海洋海滩散步时,我会把它挂在挂绳上戴在脖子上。老实说这个价格很离谱,因为硬件基本上就是一个带按钮的蓝牙麦克风,然后他们还要收转录的月费。你花 180 美元到底是为了买什么?
不管怎样我还是买了。对大多数人来说,手机应用可以做同样的事情,但我想要录制几个小时并将输出发送到我自己的流程中。这个专用的物件消除了选择。我戴上它,按住按钮,感觉到它震动,然后开始散步。
在漫长的一天里,我可以走 40,000 到 50,000 步,因为我忙着说话。出门前我不需要一个主题。我可以花前十分钟抱怨我没话可说,重复自己的话,走错路,最终找到我真正关心的东西。我继续走,因为我还没想完那个念头。
录音质量不必很棒。风、交通和半截句子过去曾让这种归档变得毫无意义。现代转录技术非常惊人。它可以恢复几乎听不见的词,LLM 可以重建一个语无伦次的论点。
人们也严重高估了人类一生会产生多少文本。几个小时的语音听起来像是一个荒谬的数据问题,但其实没那么多 token。很容易管理。
这就是为什么我愿意漫无边际地胡扯。连贯性可以稍后再说。模型可以恢复论点,保留那些听起来仍然像我的短语,向我展示我一直在兜圈子的地方,并询问它的重建是否正确。我唯一必须现场做的就是在这个念头消失前捕捉它。
Andrej Karpathy 最近描述了这个模式的十分钟桌面版:切换到语音,不加过滤地胡扯,让模型重建这个念头。他写道:“LLM 不知何故非常擅长重建冗长且不连贯的胡言乱语。”
https://x.com/karpathy/status/2079610838143623371?s=20
Plaud 就是同样的动作,只是挂在挂绳上,并且允许运行一个下午。

自 4 月 14 日以来,NotePin 在 41 天内捕获了 61 份转录稿和 586,192 个单词。这些散步现在滋养了文章、语言课程、决策、原型以及我在说话时无法预料的搜索。
你会惊讶于你在日常生活中发出了多少信号。大部分信号都消失了,因为它们没有以可发布的顺序到达。散步和通勤过去是死时间。现在它们在让一天更愉快的同时,也产生了原始材料。
流程有时会滞后。我不在乎。转录稿可以下周再总结,或者明年由更好的模型重新处理。唯一的永久性失败是没有录下来。
## OBS + 计时器
有一件事我一直很纠结,就是开始远程工作。你可以称之为拖延或 ADHD,但我不觉得这两个标签很有帮助。它们听起来像是一种道德败坏,并没有告诉我该改变什么。真正的问题通常是开始前的十分钟。一旦我开始工作,我就没事了。
我打开 OBS,这个免费的流媒体应用,录制我的屏幕,角落里有一个小框显示我的大头。然后我设置一个物理计时器,通常是 45 分钟。这是我工作能力的支柱。
我基本上是在假装直播。我解释我要做什么,一边做一边叙述,并把任何让我困惑的事情说出来。观众是我的未来自己,加上最终会阅读转录稿的任何代理。我喜欢胡扯,所以这让工作比坐在电脑前试图保持纪律、一声不吭要有趣得多。
我也会发出奇怪的声音,偶尔唱唱歌,来娱乐或尴尬一下我的未来自己。一旦你克服了这种没有任何观众的小型直播有多么奇怪的感觉,我无法想象没有它怎么工作。

录制还解决了另一个我没有好名字的问题:工作往往感觉不算数。你可能花了一个小时尝试某事,毫无进展,合上笔记本电脑,感觉这小时消失了。如果我录下来,这个会话就存在。即使是一个糟糕的片段也显示了我尝试了什么,我在哪里卡住了,我当时在想什么。
文件会上传到 Dropbox 和一个私有的、不公开的 YouTube 播放列表。YouTube 是免费的视频存储。我家有千兆上传,存储很便宜,我的协议是录制、修剪两端、发布。如果我打开 Premiere、DaVinci 或 CapCut,我就已经搞砸了。编辑将工作记录变成了内容项目。
计时器很重要,但我根据工作以不同的方式使用它。对于执行工作,我保持倒计时可见。对于开放式的思考,我会隐藏它;看着分钟流逝会让生成式工作变得更糟。无论哪种方式,计时器将边界外化,这样我就不必不断问自己是否工作了足够长的时间。
这对 ADHD(多动症)患者来说绝对是一个生活黑客。它对自学、审计时间花在哪里以及长时间保持一条线索直到完成非常有用。大声说话让注意力有一个去处。
OBS 是现实接触。
我是从观看 George Hotz 进行八小时直播并完全专注中获得这个想法的。我的私人播放列表上现在有超过 1,500 小时的录像。一个朋友曾经怀疑我是否真的在工作,所以我发给他链接。他看了一眼,立即投资了 10,000 美元。1,500 小时的 footage 是很难伪造的。
我过去常通过说未来的模型可能会用到它来证明录制一切的合理性。那个未来已经开始了。Cortex 包含了 542 个会话的转录稿,代理可以检索我在工作时说的话,以及我的聊天、散步、日记和日志。我不是为了抽象的希望有一天变得有用而保存视频。542 个会话已经是记忆了。
## Raycast 片段
Raycast 基本上就是 Mac 上普通人的命令行。我告诉人们的第一件事就是禁用 Spotlight,把 Command+Space 给 Raycast。
真正的区别在于可编程性。Spotlight 搜索 Apple 决定它应该搜索的东西。Raycast 可以从同一个框中运行扩展、脚本、片段、窗口命令和我自己的 AI 命令。我有一个扩展,输入 money 就能看到我的银行余额。如果我想要一个新的行为,我可以写一次,然后从计算机的任何地方召唤它。这感觉更像是给自己添加一个命令,而不是安装另一个应用。
片段是最小的部分,也许也是我最喜欢的。我在任何文本框中输入一个简短的触发器,Raycast 就会将其展开为时间戳、模板,或者一种与 LLM 交互的全新方式。
我最常用的是 `!bear`,清晰熊:
> 首先深入思考,回顾然后采访我,一次一个问题,向我提出经过深思熟虑的选项和建议,以提高你的清晰度,直到你达到 100% 清晰。对于每个问题,尝试指出困惑的来源可能是什么,或者回答这个问题将如何确定你需要的东西(简洁)。当足够清楚时,尝试选择多选。
当我不知道该提示什么,因为我还不知道我在想什么时,我会使用它。模型开始一次采访我一个问题。多选是诀窍。它在 95% 的情况下都能把选项搞对,从混乱的状态中选择或拒绝一个选项比从头发明一个清晰的提示要容易得多。我用它来处理卡住的项目、规格说明、大脑转储、心理状态调试以及别人的两难困境。
`!time` 甚至更简单。它展开为带有时间戳的 Markdown 行。我将其与 Raycast Notes 和按住听写配对,这比构建一些复杂的 Notion 集成更快、更可靠。本地和即时胜过集成和网络化。
清晰熊曾经是一个产品。我围绕大脑转储 -> 采访 -> 清晰度表构建了一个完整的应用程序。该产品死了。有用的部分作为击键后的六十个单词存活了下来。那段确切的文字自那以后出现在了 69 个文件中,每次都带着同样的化石般的拼写错误。我从不重打它。

## 晨间咒语 + goals.md
晨间咒语是我在工作前大声朗读的文本文件。我现在是第八个版本。称之为咒语听起来很精神,但我把它当作软件:行被添加,被审查,当它们不再有用时被删除。我最近删除了一些,因为它们只是伪装成智慧的焦虑。
它的存在是因为我可以在晚上完美地理解某件事,但第二天早上醒来时却无法访问那种感觉。如果我知道一个洞察会衰减,我会把指令放在明天的那个我实际上会看到的地方。最后一行告诉我开始计时。
goals.md 在代理界面中保持可见。它包含我目前试图达到的两个结果和现在重要的工作。代理可以探索如何到达那里,但我不希望目标被最近发生的任何对话重写。
## 有氧运动中的 GPT Live
去他妈的:GPT Live 是我用过的最好的语言学习应用。一次好的会话后,我的反应是,为什么我不每天在做有氧运动时安排三十分钟的这件事?
语言应用优化练习。我想知道我是否能用法语解释元认知,争论商业想法,角色扮演奇怪的情况,或者与虚拟女招待调情。这是第一个让我想谈论任何东西而不是完成下一课的工具。
明显的替代方案是与人交谈,但这需要安排日程、信心,以及足够的精力在公开场合把一门语言说得不好。GPT Live 有无限的耐心。我可以疲惫地到达,混合语言,在中途改变话题,胡扯直到句子出现。它仍然理解我在追求什么。
重要的功能是打断。当我不知道一个词,或者不记得某样东西是阳性还是阴性时,我立刻询问并在对话内得到答案。纠正恰好出现在不确定的时刻。抽认卡应用做不到这一点,课程通常也不会。
我在有氧运动期间运行会话,所以它在一天中有固定的位置。身体活动已经在发生;对话让它变得足够有趣,让我停止看时钟。
会话不局限于一种语言。它们大约 98% 是法语,中间会跳入粤语或韩语。如果我想练习一个韩语句子,模型会用法语解释,给我完整的韩语,把它分成片段,然后在我重复每一个时等待。法语正在成为认知的媒介,而不是被研究的对象。
我们排练我真的需要进行的对话。我可以在有氧运动进行到一半时出现,疲惫地混合三种语言,并解决我想向真人解释的困难事情。GPT Live 在语言变化时跟踪意义。
我在 2024 年已经在做一个更粗糙的版本,用语音模式“逐句、逐词”阅读母语文章。当把它与有氧运动结合起来时,它变得一致了。现在的会话读取并输入韩国器官的 _state.md,所以实时练习和结构化课程共享我知道什么以及检索失败之处的记录。

模型说出韩语句子。然后是第一个片段。我重复它。它添加下一个片段。我们继续直到我能毫无帮助地说出整个句子。
## 器官:让状态累积
上述工具仍然需要我按下录制或开始对话。器官是代理内部保持可用的代码和文件。它们保留足够多的历史,这样我就不必每次都重新解释领域。
## `organs log`
我写日记好几年了,但写日记仍然是一个你必须决定去做的活动。大多数普通事件都不足以让我打开一个单独的应用并写一篇反思性条目。
我在我已经工作的终端输入 `organs log` 后面跟一句话。该命令添加一个时间戳并将句子附加到今天的 Markdown 文件中。每天一个文件,没有表格,没有要记住的类别。过去 121 天中有大约 112 天被记录了。
如果我写我打了匹克球并加入了一个接龙游戏,代理会检查有氧运动并记录一个社交事件。我不打开习惯追踪器,然后是社交追踪器,然后是日记。我描述一次发生的事情。

原来的句子保持逐字记录。代理喜欢将具体的想法变成平淡的摘要,所以规则是在不替换来源的情况下添加解释。一个奇怪的短语、重复的抱怨或确切的思想顺序可能在后来是有用的部分。存储很便宜。重新创造句子是不可能的。
几个月后,我可以问上一次我有这种感觉时我在做什么,或者一个新想法是否真的新颖。答案带着日期和我自己的字句回来了。
## 习惯
我试过普通的习惯应用。问题是习惯追踪器本身变成了另一个习惯。如果我必须记住在它提醒我做什么之前打开一个单独的仪表板,它就已经失败了。
我的追踪器是一个直接在代理界面中渲染的 SQLite 数据库。Ctrl+H 在我的目标和当前关注点下打开它。如果我不知道下一步做什么,我按下快捷键并看看还有什么未完成。
代理看到相同的列表。如果我记录一场匹克球比赛,它可以自动检查有氧运动。如果几个习惯仍然开放,它可以从我已经使用的界面建议一个。我越使用编程代理,我就越经常看到我的习惯。这就是为什么这个追踪器在其他的失败时成功了。

我不是为了截图而表演完美的连胜。我希望今天感觉与我选择的更大的系统相连。如果一个习惯变得势不可挡,我可以缩小它的规模或跳过它而不失去记录。
## 聊天归档和 Cortex
AI 对话默认情况下奇怪地是一次性的。我可以花三个小时和 Claude 开发一个想法,然后下周打开一个新的聊天,这些工作都不可用了。
现在每个对话都被导出为 Markdown 并被总结。我有 3,520 个摘要。在一次运行中,代理在大约两个小时内处理了 3,273 个对话,提取了 77 个具有最强写作材料的对话,并找到了 101 个可能的草稿。
聊天窗口是一个界面。它不应该决定思想能存活多久。
Cortex 在一个 1.77 GB 的 SQLite 文件中索引了 9,480 个文档和 189,703 个块:每日日志、十四年的日记、对话、会议、OBS 会话和散步。它一起运行关键字和向量检索,然后重新排序结果。当语料库是你的整个人生并且你不记得你使用的词时,grep 就不是很好了。
那是原始问题。我记得有过一个想法,但不记得什么时候、在哪里或我是怎么措辞的。第一次我针对我自己多年的写作提问并得到带有日期证据的答案时,老实说有点震惊。一个模糊的直觉突然有了历史。
在做决定之前,我可以询问上一次我做出类似决定的时间。对于语言练习,代理可以检查我实际谈论的内容并从我的痴迷中生成词汇。我不必坐在那里从记忆中召唤自传式的词汇表。课程已经存在于记录中。
但搜索本身并不是那么有趣。每个人都有搜索。变得有用的是能够对我的过去运行操作。
代理可以按时间顺序遍历每个日志、散步、聊天和日记,并将想法提取到一个模式中。另一支舰队可以从这些对象构建原型。测试代理可以执行它们。我的反馈可以改变评估标准,锦标赛可以排名什么存活下来。

我还可以要求几个代理使用关于我维持、放弃、享受和返回什么的证据来模拟可能的未来。这不是预言。这是一个更好的思想实验,因为代理可以阅读比我在工作记忆中能容纳的更多的我的历史。
这就是为什么我认为“记忆”对它来说太小了。记录可以被过滤、比较、转化为模式、在代理之间拆分、用于构建原型,并根据我后来的反馈进行评分。我已经花费在思考上的计算没有被困在我做这件事的那一刻。
我曾经称 Cortex 为“不是第二大脑,而是一千个大脑”。第二大脑帮助你找到你存储的东西。Cortex 让许多思想可以加工你经历过的事情。
我相信归档是为了证据,而不是意图。它可以告诉我我说的什么以及我多久返回一次它。它不能决定我现在想要什么。
我构建的任何新东西都可以从那个历史开始,而不是一个空的提示词。
## 其他几个器官
工作日志每天保留一个 Markdown 文件,解释代码为什么改变。Git 保留差异;它不保留我试图做什么。
邮件器官对 33,166 条消息运行书面分诊策略,并产生了 82 个行动项。该策略还包含一个硬性规则,即我发送的任何内容都不能被删除。
金融器官直接与 Mercury API 对话——@mercury Personal 有一个很棒的 API,shoutout——并将交易追溯回产生它们的事件。我不使用 MCP,因为它是一个我已经可以调用的 API 的包装器。
任务保留在 Things 3 中。它有一个很好的手机界面,它的模式十年来几乎没有变化,things-cli 可以读取 SQLite 数据库。当我有一个值得替换它的手机界面时,我会替换它。
## Taskbox、Git 和 Linear
一个 Claude Code 会话几乎用不完一个 Max 订阅。如果我手动驱动它,我就会留下大部分我付了费但未使用的部分。Taskbox 是我将工作分散到单独的代理运行中而不迷失哪个提示产生了哪个结果的方式。它已经处理了 2,271 次运行和超过 3 亿个 token。
Git 几乎荒谬地适合代理工作。每个代理获得一个隔离的分支或工作树,更改真实文件,并返回一个差异。如果三个代理尝试三种方法,我可以比较代码而不是比较聊天中三个自信的段落。
拉取请求是分支必须变成一个被接受版本的地方。测试运行,审查代理检查差异,我合并存活的东西。这基本上就是 Unix 的用途。

Linear 在代理工作时持有未解决的问题。我对此有一个强烈的观点:Linear 是用于收敛状态,而不是从一个模糊的任务列表中勾选事情。一个好的问题是竞争分支、证据和审查最终定居为一个答案的家。
一旦我选择了一种方法,理由就可以变成测试、示例或指令。这就是那一刻的品味如何改变数百次后来的运行。
/will 给 Claude 一个我 58 个技能和我为自己写的七个协议的模型。Claude Code 已经知道如何调用 shell 或读取文件;这个命令告诉它 Will Chen 能做什么。
我在使用 Claude,而 Claude 正在假装使用 Will Chen,作为一个代理。
它识别出仍然需要人类判断、品味、对话或勇气的部分,然后交给我上下文和我之前写好的程序。代理通常能比我更可靠地检索我自己的能力。
## 系统:当日子开始感觉不同时
在这一点上,技术栈不再感觉像是一个工具集合。我可以合上笔记本电脑而不失去一天的状态。文件保留,代理保留它们分配的工作,下一个会话不以我试图记住一切开始。
## 晨间简报
晨间简报是我读的第一样东西。它汇集了我的目标、当前关注点、开放的习惯、任务、日历、身体数据、财务以及代理隔夜完成的任何事情。
我不打开六个应用程序并在脑海中协调它们。编程代理和我从同一个屏幕开始。如果我问今天什么重要,它可以从实际状态回答,而不是生成通用建议。
这不是我留待以后的报告。它告诉我下一步做什么。
## 在技术栈中的一天
早上我阅读简报和咒语,然后开始计时。
在白天,事件进入单行日志。散步产生转录稿。Ctrl+H 显示下一个习惯。OBS 录制成为归档的一部分。我正常使用每样东西,记录在它后面累积。
在晚上,我审查发生的事情,更新工作日志,并发送不需要我的工作。代理处理转录稿,尝试实现,运行测试,或准备选项。当我第二天早上回来时,通常有具体的东西等待审查。
## 循环
对我来说,代理是一个循环。它读取一个状态,做某事,检查结果,然后继续或停止。把每个子代理当作一个有性格的单独员工对待只会让协调更难推理。
散步变成转录稿。Cortex 检索相关材料。Taskbox 对其运行一个操作。结果作为文件、分支、日志条目或对另一个器官的更新回来。如果一个代理弄错了什么,纠正会在下次运行之前变成测试或指令。

每个有用的循环都接触到模型之外的东西。完成的锻炼更新追踪器。银行交易更新财务状态。测试说代码是否有效。没有那些答案,代理只会变得更擅长同意自己。
模型读取文件,调用命令,并在域之间移动信息。我可以替换模型,因为仓库、记录、策略、测试和权限在会话中存活下来。

## 从使用 AI 到让 AI 系统运行
我没有决定构建第二大脑或生活操作系统。我在真实的一天中继续使用编程代理并改变任何让我烦恼的东西。因为这些改变住在一个仓库里,每个新修复都从上一个开始。
这是我认为人们在个人 AI 上错过的部分。架构可以稍后出现。第一个要求是让纠正保留的地方。把追踪器放在你已经工作的地方。如果录制让它更容易开始,就录制会话。如果语言练习变得令人愉快,就把它移到有氧运动中。当修复有效时,保留它。
大多数 AI 演示让生成看起来像是突破。我们已经生成了太多的东西。对我来说更大的变化是连接变得廉价:散步可以滋养文章,日志条目可以更新习惯,旧的对话可以变成原型,一次纠正可以到达每一个后来的代理。
2024 年,我写道:“尽可能多地记录数据——我相信未来的计算会更便宜,并允许我们复习我们采取的所有笔记。”当时,这是一个赌注。现在记录可以被搜索、转换、分布在代理中、测试并作为行动返回。
这篇文章来自同样的技术栈。我散步时口述了部分内容,检索了相关的日志和对话,将部分通过单独的代理发送,根据日期证据检查主张,并在同一个仓库中不断修改结果。
AI Organs 只有几个月大,但它已经稳定了我生活中和精神能量的惊人数量。洞察有一个去处。它可以变成命令、策略、测试、课程、草稿或另一个代理的任务,而不是变成另一件我应该记住的事情。
这就是转变。显然,我仍然使用 AI。但越来越多的时候,我有 AI 系统在运行。