翻译腔 「多个 agent 一起尝试复现一个被报告的问题」硬套被动。「他们能干又蠢,而且很可教」照搬英文结构。引文里(across many agents working on 10 different problems)只译成「在 10 个不同问题上切换上下文」,丢了「很多 agent」。好的地方是「到处去点东西,感觉近乎野蛮」「compaction 快得离谱」读起来很自然。
gemini-3.1-pro:第七名。逐词直译最多,翻译腔最重,还改坏了一处结构。
翻译腔 「但我不能说谎。」直译(I can't lie though)。还有「天哪,它能写而且会写出一大堆代码」「作为一个 cc 用户」「agent 会谄媚地不惜一切代价」「这完全就像管理人类一样!」,以及「瓶颈依然在于我要努力记住,并在我自己的上下文里保持我的 agent 都在做些什么」。
误译和术语 原文(The sticking point for me)译成「对我来说,难点在于」,是误译。「你把产品做得很 hackable」是中英夹杂。「带着从过去经验里学到的技能」违反术语表(skills 应保留)。「大部分快捷键还在我的上下文里」丢了原文「上下文窗口」那个玩笑。
I need to get something off my chest. Before my interview @cursor_ai, I had never actually used Cursor.
本站旧译文
有件事我想先说清楚。在去面试 @cursor_ai 之前,我其实从没真正用过 Cursor。
claude-opus-5-5
有件事我得坦白。去 @cursor_ai 面试之前,我其实从没用过 Cursor。
gpt-5.6-sol
有件事我得坦白。在参加 @cursor_ai 的面试之前,我其实从未用过 Cursor。
gemini-3.1-pro
我得坦白一件事。在 @cursor_ai 面试之前,我其实从没用过 Cursor。
grok-4.7
我得先坦白一件事。面试 @cursor_ai 之前,我从来没真正用过 Cursor。
kimi-k3
有件事我得坦白。去 @cursor_ai 面试之前,我其实从来没用过 Cursor。
glm-5p3
有件事我得先一吐为快。去 @cursor_ai 面试之前,我其实从来没真正用过 Cursor。
英文原文 · 第 3 段
At Meta, Claude Code was explosively taking off. I even paid for a personal $200 a month plan for my side projects. I loved how simple it was, and how quickly I could feel productive. The sticking point for me was in developing my own set of skills that turned cc into almost anything I wanted. I even started developing my own agent orchestrator tool on top of it.
本站旧译文
在 Meta,Claude Code 当时火得不行。我甚至为自己的 side project 付了每月 200 美元的个人方案。我喜欢它足够简单,上手就能感到高效。卡住我的地方是:我想自己做一套 skills,把 cc 拧成几乎什么都能干的样子。我甚至开始在它上面自己写 agent 编排工具。
在 Meta 的时候,Claude Code 正火得一塌糊涂。我甚至自掏腰包订了每月 200 美元的个人套餐,用来做自己的业余项目。我喜欢它的简单,喜欢上手就能出活。最让我着迷的是打造自己的一套 skills,把 cc 变成几乎任何我想要的样子。我甚至在它之上开始写自己的 agent 编排工具。
glm-5p3
在 Meta 的时候,Claude Code 正在爆发式流行。我甚至为副业项目自费买了每月 200 美元的个人套餐。我喜欢它的简单,也喜欢自己这么快就有产出感。最让我着迷的,是自己写一套 skills,把 cc 变成几乎任何我想要的东西。我甚至开始在它上面做自己的 agent 编排工具。
英文原文 · 第 4 段
During my onsite interview, I used Cursor over the course of 2 days, using it to build the interview project. This was before Cursor 3's release, so I was using the Editor Window. I've been using vscode for so many years at this point that most of the keyboard shortcuts were still in my context window, so getting back on the IDE was not too difficult. I can't lie though. During the first hour or two I was certainly missing the cli. Clicking on things felt almost barbaric. But there were a few things that really stood out to me.
First, the models I was used to at the time - Opus and Codex - felt smarter, somehow. And it was amazing being able to switch models on the fly and use both of them at the same time on different parts of my project (Opus for the frontend, Codex for systems). Prior to my interview, I was already raving about multi-model adversarial review, so being able to do this natively in the UI felt very natural. Even better was the ability to spawn subagents of different models, so I could get the best of both worlds in one conversation.
Second, compaction was insanely quick. As a cc user I was used to compacting taking many minutes, and so I was always in a state of constant vigilance of my context and plan usage. So I was utterly shocked at how fast it was in Cursor. So much so that I basically never needed to look at how much context I was using. It just worked, while I often felt like the model just got super dumb after compaction in cc.
本站旧译文
第二,compaction 快得离谱。用 cc 时,compact 往往要好几分钟,所以我总在盯着上下文和 plan 用量。Cursor 里快到我彻底震惊,以至于我几乎从不需要看还剩多少上下文。它就是能用。而在 cc 里,compact 之后我经常觉得模型突然变笨了很多。
claude-opus-5-5
第二,压缩上下文快得离谱。作为 cc 用户,我习惯了压缩一次要好几分钟,所以一直时刻盯着上下文和套餐用量。Cursor 里压缩这么快,我完全震惊了。快到我基本再也不用去看上下文用了多少。它就是好用。而在 cc 里,我常常觉得压缩之后模型一下子变得特别笨。
gpt-5.6-sol
第二,上下文压缩快得惊人。作为 cc 用户,我早已习惯压缩要等上好几分钟,所以总得时刻盯着上下文和套餐用量。Cursor 的速度快得让我无比震惊,快到我基本不用再查看用了多少上下文。它就是能正常工作;而用 cc 时,我常觉得模型压缩后一下子变得特别笨。
gemini-3.1-pro
第二,compaction 快得离谱。作为一个 cc 用户,我习惯了 compaction 动辄好几分钟,所以我总是时刻警惕我的上下文和套餐用量。所以 Cursor 里的速度彻底震惊了我。以至于我基本上不用去看用了多少上下文。它就是能用,而在 cc 里我经常觉得模型在 compaction 后就变傻了。
grok-4.7
第二,压缩快得离谱。用 cc 的时候,我习惯了压缩要花好几分钟,于是一刻不停地盯着上下文和套餐用量。所以 Cursor 里的速度让我彻底震住了。快到我基本上不用再看上下文用了多少。它就是能用。在 cc 里,我却常常觉得模型一压缩完就变得特别笨。
kimi-k3
第二,压缩快得离谱。作为 cc 用户,我习惯了压缩一次要好几分钟,所以总是时刻盯着自己的上下文和套餐用量。Cursor 快得让我震惊,快到我基本不用再看自己用了多少上下文。它就是能用。而在 cc 里,我常常觉得模型一压缩完就变蠢了一大截。
glm-5p3
第二,压缩快得离谱。作为 cc 用户,我习惯了压缩要花好几分钟,所以一直对自己的上下文和 plan 用量保持警惕。Cursor 里的速度让我彻底震惊,快到我基本不用看自己用了多少上下文。它就是能用,而在 cc 里,压缩之后我常觉得模型直接变傻了。
英文原文 · 第 7 段
And the third thing I noticed was how much GUIs could offer over TUIs. Being able to open your app directly in Cursor's browser and make design changes with Design Mode felt intuitive and got me thinking about how much purpose-built UIs could make agentic coding more effective.
Since joining at the end of March, I've been working primarily on Cursor 3's Agent Window, and using it as my daily driver. While I still think cc is a cool product with a great team, I have noticed that its simplicity tends to drive people to want to build their own abstractions wrapping it. In my last job, it felt like there was a new internal orchestrator tool built on top of cc announced every week.
本站旧译文
三月底入职以来,我主要做 Cursor 3 的 Agent Window,也把它当日常主力。我仍然觉得 cc 是个很酷的产品,团队也很棒,但我注意到它的简单往往会推着人去包一层自己的抽象。上一份工作里,感觉每周都会冒出一个基于 cc 的内部编排工具。
claude-opus-5-5
三月底入职以来,我主要在做 Cursor 3 的 Agent Window,也把它当作日常主力工具。我仍然觉得 cc 是个很酷的产品,团队也很棒。但我注意到,它太简单了,往往让人想在外面再包一层自己的抽象。在上一份工作里,感觉每周都有人宣布一个基于 cc 的新内部编排工具。
gpt-5.6-sol
自从三月底加入以来,我主要负责 Cursor 3 的 Agent Window,也一直把它作为日常主力工具。虽然我仍然认为 cc 是一款很酷的产品,背后也有一支出色的团队,但我发现,它的简单往往会让人想在外面再包一层,构建自己的抽象。在我上一份工作里,感觉每周都会有人发布一个基于 cc 构建的新内部编排工具。
gemini-3.1-pro
自 3 月底加入以来,我主要在做 Cursor 3 的 Agent Window,并把它作为日常主力。虽然我仍然觉得 cc 是个很酷的产品,团队也很棒,但我注意到它的简单往往会驱使人们想在它外面包装自己的一套抽象。在上一份工作里,感觉每周都会宣布一个构建在 cc 之上的新内部编排工具。
grok-4.7
三月底加入以来,我主要在做 Cursor 3 的 Agent 窗口,并且把它当成我每天的主力工具。我仍然觉得 cc 是个很酷的产品,团队也很强。但它的简单,往往会让人想自己做一层抽象,把它包起来。上一份工作里,感觉几乎每周都会宣布一个新的内部编排工具,就建在 cc 上面。
kimi-k3
三月底入职以来,我主要在做 Cursor 3 的 Agent Window,日常也拿它当主力工具。我依然觉得 cc 是个很酷的产品,团队也很棒,但我注意到,它的简单反而会推着人想自己在它外面包一层抽象。在上一份工作时,感觉每周都有一个建在 cc 之上的内部编排工具发布。
glm-5p3
三月底入职以来,我的工作主要在 Cursor 3 的 Agent Window 上,它也是我每天的主力。我仍然觉得 cc 是个很酷的产品,团队也很棒,但我注意到,它的简单反而会推着大家想在它外面包一层自己的抽象。在上一份工作里,感觉每周都有人宣布一个新的、基于 cc 的内部编排工具。
英文原文 · 第 10 段
@bcherny talks a lot about this idea of "latent demand":
本站旧译文
@bcherny 经常谈「latent demand」这个概念:
claude-opus-5-5
@bcherny 经常讲「潜在需求」这个概念:
gpt-5.6-sol
@bcherny 经常谈到「潜在需求」这个概念:
gemini-3.1-pro
@bcherny 经常谈到「潜在需求」这个概念:
grok-4.7
@bcherny 经常讲「潜在需求」这个想法:
kimi-k3
@bcherny 经常讲「潜在需求」这个说法:
glm-5p3
@bcherny 经常谈「潜在需求」这个想法:
英文原文 · 第 11 段
"There's this really old idea in product called latent demand... you build a product in a way that is hackable, that is kind of open-ended enough that people can abuse it for other use cases. Then you see how people abuse it and then you build for that."
本站旧译文
"There's this really old idea in product called latent demand... you build a product in a way that is hackable, that is kind of open-ended enough that people can abuse it for other use cases. Then you see how people abuse it and then you build for that."
This was exactly it! Folks converging on orchestration tools exposes the latent demand in that using a cli just makes you, the human, the orchestrator.
But every agent workflow I had used was focused on the wrong thing. Running multiple CLIs in a GUI was missing the point entirely. The approach I was interested in was building trust in agents.
As a former engineering manager, I quickly realized that managing agents felt similar to building a human engineering team. New hires need to be onboarded so they understand the codebase, but also how work gets done. They join already pre-trained with skills they acquired from their past experiences: how to debug, how to write high quality code and tests, and how to communicate, to name a few.
Agents are like new hires in a constant state of amnesia and idiocy. They don't remember what you tell them, and they never really learn anything new. But we can equip them with rules, skills, tools, and long term memory which can approximate that. They're capable yet stupid, and very teachable. And I saw their failure modes as opportunities to teach them everything I know about doing deep, rigorous engineering.
Because when there is no rigor, agents will sycophantically do whatever it takes to write that code you asked for. And boy, can and will it write a lot of it. Naive parallelization just makes them write slop faster.
i'm increasingly convinced that the value of orchestrating many agents in parallel comes from going deep, not broad. you want to go deeper on a single or a few problems, so you can maximize your chances of getting great results:
best of N style races to find the best solution
adversarial review
multiple agents trying to repro a reported issue
use different models for different types of workloads
to quote @mattpocockuk, "code is not cheap".
this beats trying to context switch across many agents working on 10 different problems. the bottleneck is still me trying to remember and keep in my own context window what my agents work on. i think there will be ways to solve this with better ui.
there's also the element of trust. the more trust you have, the higher up the perspective ladder you can go. when you have little trust, you must micromanage each and every agent. when you can let go, you can delegate more. this is exactly like managing people!
I do think that agent orchestration can be done productively. But we need to go depth first.
本站旧译文
我确实认为 agent 编排可以做得有产出。但我们需要 depth first。
claude-opus-5-5
我确实认为 agent 编排可以做得高效。但我们得深度优先。
gpt-5.6-sol
我确实认为,agent 编排可以富有成效。但我们得先向深处走。
gemini-3.1-pro
我确实认为 agent 编排可以做得富有成效。但我们需要优先做深。
grok-4.7
我确实觉得,agent 编排可以做得很有成效。但我们得先做深。
kimi-k3
我确实认为 agent 编排是可以做出成效的。但我们得先往深里走。
glm-5p3
我确实认为 agent 编排可以高效地做出来。但得先往深里做。
英文原文 · 第 20 段
I'm open sourcing pstack, my personal set of skills and engineering principles that I use everyday to build @cursor_ai. I started developing early iterations of these skills in my side projects, and have been refining them ever since.
[图片:Cursor's company leaderboard. My skills were used 9k times this week!]
本站旧译文
[图片:Cursor 的公司排行榜。我的 skills 这周被用了 9k 次!]
claude-opus-5-5
[图片:Cursor 公司内部排行榜。这周大家用了我的 skills 9000 次!]
gpt-5.6-sol
[图片:Cursor 公司内部排行榜。我的 skills 本周使用了 9 千次!]
gemini-3.1-pro
[图片:Cursor 公司排行榜。这周我的 skills 被用了 9 千次!]
grok-4.7
[图片:Cursor 的公司排行榜。我的 skills 这周使用了 9k 次!]
kimi-k3
[图片:Cursor 公司内部排行榜。我的 skills 这周被用了 9k 次!]
glm-5p3
[图片:Cursor 公司内部排行榜。这周我的 skills 被用了 9k 次!]
英文原文 · 第 25 段
Cursor's company leaderboard. My skills were used 9k times this week!
本站旧译文
Cursor 的公司排行榜。我的 skills 这周被用了 9k 次!
claude-opus-5-5
Cursor 公司内部排行榜。这周大家用了我的 skills 9000 次!
gpt-5.6-sol
Cursor 公司内部排行榜。我的 skills 本周使用了 9 千次!
gemini-3.1-pro
Cursor 公司排行榜。这周我的 skills 被用了 9 千次!
grok-4.7
Cursor 的公司排行榜。我的 skills 这周使用了 9k 次!
kimi-k3
Cursor 公司内部排行榜。我的 skills 这周被用了 9k 次!
glm-5p3
Cursor 公司内部排行榜。这周我的 skills 被用了 9k 次!
英文原文 · 第 26 段
pstack teaches agents to be more rigorous using multiple models. I've taken all the failure modes I've observed and turned them into skills. The heart of the plugin is /poteto-mode, which is a higher order skill that gives agents the right playbook to follow for a given task. The goal is not maximal LOC, but the opposite: maximum impact with the least amount of code.
I need to get something off my chest. Before my interview @cursor_ai, I had never actually used Cursor.
At Meta, Claude Code was explosively taking off. I even paid for a personal $200 a month plan for my side projects. I loved how simple it was, and how quickly I could feel productive. The sticking point for me was in developing my own set of skills that turned cc into almost anything I wanted. I even started developing my own agent orchestrator tool on top of it.
During my onsite interview, I used Cursor over the course of 2 days, using it to build the interview project. This was before Cursor 3's release, so I was using the Editor Window. I've been using vscode for so many years at this point that most of the keyboard shortcuts were still in my context window, so getting back on the IDE was not too difficult. I can't lie though. During the first hour or two I was certainly missing the cli. Clicking on things felt almost barbaric. But there were a few things that really stood out to me.
First, the models I was used to at the time - Opus and Codex - felt smarter, somehow. And it was amazing being able to switch models on the fly and use both of them at the same time on different parts of my project (Opus for the frontend, Codex for systems). Prior to my interview, I was already raving about multi-model adversarial review, so being able to do this natively in the UI felt very natural. Even better was the ability to spawn subagents of different models, so I could get the best of both worlds in one conversation.
Second, compaction was insanely quick. As a cc user I was used to compacting taking many minutes, and so I was always in a state of constant vigilance of my context and plan usage. So I was utterly shocked at how fast it was in Cursor. So much so that I basically never needed to look at how much context I was using. It just worked, while I often felt like the model just got super dumb after compaction in cc.
And the third thing I noticed was how much GUIs could offer over TUIs. Being able to open your app directly in Cursor's browser and make design changes with Design Mode felt intuitive and got me thinking about how much purpose-built UIs could make agentic coding more effective.
Building Cursor with Cursor
Since joining at the end of March, I've been working primarily on Cursor 3's Agent Window, and using it as my daily driver. While I still think cc is a cool product with a great team, I have noticed that its simplicity tends to drive people to want to build their own abstractions wrapping it. In my last job, it felt like there was a new internal orchestrator tool built on top of cc announced every week.
@bcherny talks a lot about this idea of "latent demand":
"There's this really old idea in product called latent demand... you build a product in a way that is hackable, that is kind of open-ended enough that people can abuse it for other use cases. Then you see how people abuse it and then you build for that."
This was exactly it! Folks converging on orchestration tools exposes the latent demand in that using a cli just makes you, the human, the orchestrator.
But every agent workflow I had used was focused on the wrong thing. Running multiple CLIs in a GUI was missing the point entirely. The approach I was interested in was building trust in agents.
As a former engineering manager, I quickly realized that managing agents felt similar to building a human engineering team. New hires need to be onboarded so they understand the codebase, but also how work gets done. They join already pre-trained with skills they acquired from their past experiences: how to debug, how to write high quality code and tests, and how to communicate, to name a few.
Agents are like new hires in a constant state of amnesia and idiocy. They don't remember what you tell them, and they never really learn anything new. But we can equip them with rules, skills, tools, and long term memory which can approximate that. They're capable yet stupid, and very teachable. And I saw their failure modes as opportunities to teach them everything I know about doing deep, rigorous engineering.
Because when there is no rigor, agents will sycophantically do whatever it takes to write that code you asked for. And boy, can and will it write a lot of it. Naive parallelization just makes them write slop faster.
i'm increasingly convinced that the value of orchestrating many agents in parallel comes from going deep, not broad. you want to go deeper on a single or a few problems, so you can maximize your chances of getting great results:
best of N style races to find the best solution
adversarial review
multiple agents trying to repro a reported issue
use different models for different types of workloads
to quote @mattpocockuk, "code is not cheap".
this beats trying to context switch across many agents working on 10 different problems. the bottleneck is still me trying to remember and keep in my own context window what my agents work on. i think there will be ways to solve this with better ui.
there's also the element of trust. the more trust you have, the higher up the perspective ladder you can go. when you have little trust, you must micromanage each and every agent. when you can let go, you can delegate more. this is exactly like managing people!
I do think that agent orchestration can be done productively. But we need to go depth first.
I'm open sourcing pstack, my personal set of skills and engineering principles that I use everyday to build @cursor_ai. I started developing early iterations of these skills in my side projects, and have been refining them ever since.
These skills have become some of the most used skills by the Cursor team, so I'm excited to share it with all of you.
[图片:Cursor's company leaderboard. My skills were used 9k times this week!]
Cursor's company leaderboard. My skills were used 9k times this week!
pstack teaches agents to be more rigorous using multiple models. I've taken all the failure modes I've observed and turned them into skills. The heart of the plugin is /poteto-mode, which is a higher order skill that gives agents the right playbook to follow for a given task. The goal is not maximal LOC, but the opposite: maximum impact with the least amount of code.
本站旧译文
[图片:我怎么用 Cursor]
有件事我想先说清楚。在去面试 @cursor_ai 之前,我其实从没真正用过 Cursor。
在 Meta,Claude Code 当时火得不行。我甚至为自己的 side project 付了每月 200 美元的个人方案。我喜欢它足够简单,上手就能感到高效。卡住我的地方是:我想自己做一套 skills,把 cc 拧成几乎什么都能干的样子。我甚至开始在它上面自己写 agent 编排工具。
三月底入职以来,我主要做 Cursor 3 的 Agent Window,也把它当日常主力。我仍然觉得 cc 是个很酷的产品,团队也很棒,但我注意到它的简单往往会推着人去包一层自己的抽象。上一份工作里,感觉每周都会冒出一个基于 cc 的内部编排工具。
@bcherny 经常谈「latent demand」这个概念:
"There's this really old idea in product called latent demand... you build a product in a way that is hackable, that is kind of open-ended enough that people can abuse it for other use cases. Then you see how people abuse it and then you build for that."
自从三月底加入以来,我主要负责 Cursor 3 的 Agent Window,也一直把它作为日常主力工具。虽然我仍然认为 cc 是一款很酷的产品,背后也有一支出色的团队,但我发现,它的简单往往会让人想在外面再包一层,构建自己的抽象。在我上一份工作里,感觉每周都会有人发布一个基于 cc 构建的新内部编排工具。