Lessons

What Not to Reinvent

What Not to Reinvent

There's a trap a lot of AI startups fall into right now: because the technology is new, we assume the whole product has to feel new too. New interface, new language, new workflow, new category. We've been guilty of it at Earmark.

We've been thinking about a framework Mark Pincus talks about called Proven / Better / New, and it gave me a sharper way to hold this. Some parts of a product should be proven: familiar behaviors customers already understand and expect. Some parts should be better: faster, cheaper, less annoying, obviously superior to how people do it today. And a small number of parts should be new - bets that might change how people work, and probably won't.

What makes the framework useful isn't the three buckets. It's the different standard of evidence each one demands.

Proven means you don't get to have an opinion. Pincus's phrasing was blunt: whatever you're not innovating on, you copy, and you don't even question it. Not because copying is virtuous, but because judgment is the scarcest thing a startup has and spending it on solved problems is how you die tired. Meetings should be captured automatically. You should be able to find an old conversation, search it, share it, and see what happened. Summaries, action items, history, permissions. Fathom, Otter and the rest have taught the market what to expect. Inventing a new mental model for any of it would make Earmark worse, not differentiated. Those parts should feel boring.

Better has an unusually high bar, and I don't think most founders apply it honestly. The test is that ten out of ten existing users would say it's better. Not "some segment prefers it." Not "it's better once you understand why." Ten out of ten. Free is better. Half price is better. Faster is better. Fewer steps is better. If a reasonable current user could look at your change and say "I actually liked it the old way," it isn't better - it's a bet you've misfiled.

Which brings me to the part of this essay I found uncomfortable to write.

Auditing our own buckets

The clearest "better" we have is what happens after the meeting. The normal AI workflow today is: finish a conversation, open the transcript, paste context into Claude or ChatGPT, explain what you want, generate a ticket or a status update or a follow-up, edit the output, move it into another system, repeat next week. AI made each step easier and in the process invented a new kind of busywork - we now spend real time assembling context for models that weren't in the room when the work was discussed.

Removing that passes the ten-out-of-ten test. If you discussed a feature, the requirements should already exist. If the team made a decision, it's captured. If engineering needs tickets, you shouldn't have to reconstruct the conversation for a model afterward. Nobody prefers doing the reassembly by hand. The meeting should produce the work, not another document you have to process.

But I've been putting other things in the "better" bucket that don't survive the same test.

Botless capture is the honest example. We built it because we think a bot joining the call creates social and operational friction, and I believe that. But "no bot in the meeting" is not something ten out of ten users would call better on sight - some people like seeing the recorder, because it's a visible consent signal and a shared cue that notes are being taken. That's not a worse product. That's a hypothesis about friction, which makes it a new bet wearing a better costume. Retention controls for privacy-sensitive companies are similar: unambiguously better for a specific buyer, neutral-to-worse for a user who'd rather everything just be kept forever.

Misfiling a bet as an improvement is the dangerous error, because you stop testing it. Improvements get shipped and forgotten. Bets get instrumented. When I moved botless from "better" to "new" in my own head, the obvious next question became: what would tell us we're wrong about it? We didn't have an answer. We do now.

There's a clean test for sorting these, and it came out of the interview almost in passing. Someone raised near-zero-latency AI responses: is that better or new? Lower latency is better - everyone says yes to it. The feature you build with it - an AI that listens and participates live - is new. The capability is better; the behavior it enables is a bet. Run your roadmap through that split and a surprising amount of what you called "better" turns out to be behavior change you haven't earned yet.

The new bet, and what would kill it

Our bet is that conversations are becoming a context layer for software.

An enormous amount of what a company knows never reaches Jira, Salesforce, Notion, or Slack. It lives in customer calls, architecture debates, planning meetings, one-on-ones - someone explaining why a decision got made, what a customer actually meant, which tradeoff mattered, what the team is worried about. Then the meeting ends and that context becomes hard to use again.

The first generation of AI meeting products made it easier to capture. That was real progress, and I think it's also transitional. A transcript is still a document: someone has to find it, read it, interpret it, and act. The bet is that the conversation itself becomes usable context - that what an organization learns accumulates across conversations, people, and time. You should be able to ask what customers have said about a feature without remembering which five calls contained the feedback. Why a decision was made, without knowing which meeting. What's changed since the last account review, where a project is blocked, what commitments the team has made. And eventually, agents doing work on top of that same context.

Two things make me hold this loosely.

The first is Pincus's discipline about new ideas: start from the proposition that yours is probably wrong. New is what gets someone to try your product - the back of the cereal box. It's usually not why they come back. A clever feature can drive every bit of your trial volume and have nothing to do with retention. If shared organizational context turns out to be the thing people demo and never use, I want to find that out in a quarter, not in three years of roadmap.

So here is what would falsify it for us: if teams who've had multi-person, cross-meeting context available for a month are still opening individual meetings to find things, the unit of knowledge is the meeting and we're wrong. If cross-conversation queries are something people run once out of curiosity and never again, we're wrong. If the artifacts we generate get heavily rewritten before anyone sends them, we haven't understood the work - we've summarized it. Those are checkable, and none of them require a debate.

The second thing is the cost curve. Most of what makes accumulated conversational context expensive today gets cheaper on a schedule. Pincus makes the point about free as a business plan - anything that can be free will be free - and the corollary is that the right way to design right now is to ask what the product looks like when the compute is effectively unlimited. We're deliberately building toward the version that's obvious in two years and slightly wasteful today, because the alternative is building around a constraint that's about to disappear.

Passionate about the instinct, dispassionate about the version

The hardest part of this is emotional rather than analytical. Founders get attached to the novel part, because it's the most fun to build and the most exciting to explain. The version of me that gets to keep this framework useful is the one who can stay committed to the instinct while being genuinely indifferent to any particular expression of it.

So: we have strong conviction that conversations should become dramatically more useful than meeting notes, that AI should understand the work being discussed rather than summarize it afterward, and that turning a company's conversations into context for people and agents is worth building. I have no attachment to which expression wins. Maybe it's real-time intelligence during the conversation. Maybe it's that useful work is already waiting when the meeting ends. Maybe it's shared organizational context. Maybe it's agents operating on top of it. We're testing all of them and we don't need more than one to be right.

Pincus has a line about knowing when it's working - that when you've really got it, you don't have to tell anyone to work harder, because everyone can see it. When it's not quite right, it's debatable, and you find yourself hunting for one more data cut to justify continuing. That's a useful tell. If we're still arguing about whether the new thing is working, it isn't.

In practice, this means we want to be excellent at the proven job of capturing conversations, unambiguously better at turning those conversations into finished work, and ambitious - but instrumented - about the possibility that conversations become shared context for humans and agents.

AI gives us permission to reinvent almost everything. The harder decision, and I think the more valuable one, is deciding what not to.

Mark Barbir

Earmark Co-founder & CEO

Let your meetings finish the work.

Earmark turns conversations into finished work — so the follow-up is already started when the call ends.