Pingbo
English
Join the beta

How Pingbo Decides What Matters

Pingbo’s job is to learn what matters to you and keep it all in order. What it handled, what’s waiting, what still needs you. You can check in when you have time, without constant interruptions from new email.

A highlighted email request connected to a Needs you tray, beside a Waiting tray and a ruler.

Pingbo has one job, and it is harder than it sounds. It tells you what actually matters in your inbox, and leaves the rest alone.

Everything else Pingbo does rests on that call. The reply it drafts is only useful if the email really asked for one. The follow-up it schedules is only useful if you really were waiting on somebody. Get the call wrong one way and you miss the message that mattered. Get it wrong the other and you are back to checking every email yourself, no better than an ordinary inbox app. So the accuracy of that one decision is the thing we measure first and check hardest.

This article shares how we measured it, and the full results. Every number in it, the answer each model gave for each email, and the code that recomputes them are published in pingbo-bench.

Sorting mail is an old problem

The first thing anyone taught a computer to do with email was throw some of it away. In 1997 a team at Microsoft Research began building a junk-mail filter on Bayesian methods, and in 1998 published A Bayesian Approach to Filtering Junk E-Mail. The idea fits in a sentence. Count how often each word turns up in junk and how often in real mail, and when a new message arrives, multiply the odds of its words together. “Free” and “winner” pull one way, “meeting” and “agenda” the other. In 2002 Paul Graham’s essay A Plan for Spam gave programmers a version they could train on their own mailbox in an afternoon, and it spread everywhere.

Counting words could not hold the line, because spammers learned the words. Gmail moved from rules to learned models and, by 2019, to large ones fed signals well beyond the text, such as who sent the mail and from where, whether they proved who they were, what the same message did across billions of inboxes, and what people clicked when they pressed “report spam”. That is how Gmail keeps more than 99.9% of spam out. Through all of it the question stayed the same. Is this message junk?

What the filter guards has not changed either. Gmail, Outlook and Yahoo still hand you what email has always handed you, a list of conversations with the newest on top, each thread on its own. The inbox knows what you have opened. It does not know what you have handled, what you are waiting on, or what you still owe, and when one piece of work runs across five threads and three weeks, you are the one holding it together. Pingbo’s unit is not the message or the thread but the Matter, every thread on one piece of work gathered into one place, a renewal, a hire, a dispute, with its people, its commitments, its deadlines and its next move, kept until the work is done.

Pingbo asks a different question of it, and a harder one. Not whether a message is junk but whether a Matter needs you, and what it needs. No count of words answers that, because the answer is in what the sender is asking, in the sender’s own words, and in whose move it is. A word filter sees “please” and “approve” and cannot tell whether the request is aimed at you or was forwarded to you for information. A language model can read the ask. Left to itself, though, it gives back a paragraph of opinion, and a product needs a decision it can stand behind.

The decision everything depends on

That decision is made about a Matter, not a message. Gmail has already thrown the junk away before Pingbo sees a thread. Pingbo gathers the work that remains into Matters, every live Matter sits in one of two lanes, and which lane is the decision everything else follows from.

Needs you is the short list of what you owe, and the card says which kind. A To Respond means the email asks a question you should answer, and the reply is already drafted before you open it. You review and send it, reword it, or discuss it with Pingbo first. A To Act means it asks for something outside the mail, filing a form or approving an invoice, and the card names the step. By the time you look, the board says what moved overnight and the replies are waiting.

Waiting is everything still live that asks nothing of you today. You are waiting on somebody else, or the thread is worth an eye. Pingbo judges the follow-up once, sets the date, and keeps the promise. When the date arrives the card comes back as a chase, naming who to nudge and for what, taken from the thread’s own record rather than a guess. You press Follow up now, or Not yet.

Done takes a Matter off the board, and Pingbo only ever proposes it, with the messages it rests on. You accept or decline, and nothing closes by itself.

Everything that is not work gets a home too, so it never reaches the board. Screened holds the bulk, automated and reference traffic you would never act on, receipts, notifications, the list you never read, kept and searchable with nothing deleted. Reading holds the mail worth a look and nothing more, a newsletter, a status update, a notice you read once.

The board with Needs you selected, listing the two Matters that need you. A Matter in Needs you with a drafted To Respond reply and a To Act step. A Matter in Waiting whose follow-up is due, with a Follow up now button. A Waiting Matter that Pingbo proposes as Done, with Not yet and Mark done buttons. Reading, with newsletters worth a look sorted into Unread, Past and Saved. Screened, with a sign-in code, a receipt and a delivery notice judged not important.
Figure 1. Where an email lands in Pingbo: Needs you, with To Respond and To Act; Waiting, with Follow up now and Not yet; Done; and, off the board, Reading and Screened.

So the model is never asked for a paragraph of opinion. It is asked to fill in a form, one question at a time, and each answer decides what it may say next. It says whose move it is, yours or theirs, before it is allowed to pick a lane. Then it commits to one of four answers, each of which you have already met. A reply is owed, and here is the draft, which becomes a To Respond. Something is owed but not a reply, which becomes a To Act. You are waiting on somebody else, which is Waiting. Or the thread is finished, which is a Done it proposes and you accept or decline. And to claim you owe a reply, it has to quote the words in the thread that ask for one, or the claim does not stand.

That form is what turns a model into a product. It is also more places to be wrong than a yes or a no, which is why we could not take the model’s word for any of it. We had to measure.

The bar we hold it to

We score against 508 real work emails from a published research dataset, Sappelli et al. (2016), Information Sciences, joined to the Enron archive for the message text. Two annotators had read 1,145 of them and recorded what each one expected of the person who received it. We keep only the 508 where both said the same thing. If two careful readers disagree about an email, it has no business marking anyone’s homework.

They sorted each email one of four ways. 112 needed an immediate reply. 67 needed a postponed reply, an answer you owe but not today. 285 were something you are accountable for knowing and owe no answer to. 44 you could ignore. Pingbo has to know which matters need your attention, so the first two become Needs you; what you are accountable for maps to Waiting, and what you could ignore is what Screened keeps off the board in Pingbo. This evaluation scores one thing only: whether the email needs a reply. So that is 112 + 67 = 179 that owe a reply, and 285 + 44 = 329 that do not.

That ratio is the floor. Answer “no” to all 508 and you are right about the 329 and wrong about all 179, which is 329 ÷ 508 = 64.8% from a system that never read a word. Anything below that is worse than doing nothing.

Accuracy is the share it gets right, and one number hides which mistakes it made. Two systems can both score 80% here, which is 406 of 508. One catches 77 of the 179 and flags nothing extra. The other catches 160 and flags 83 you did not need. The first leaves 102 emails you owed unread. The second leaves 19. So we publish three numbers rather than one. Accuracy, the share it got right overall. Recall, how much of what you owe it catches. And precision, how much of what it shows was really owed.

We hold 32 models to that bar, chosen to answer three questions. 16 from OpenAI and 11 from Anthropic show what the frontier can do. Qwen3.6, an open-weight model we run on our own hardware, shows what Pingbo can do without renting one. And four small models show how much of this job a phone could do. Gemma and two builds of Ornith are small enough that a phone could one day carry them, and Apple’s on-device model, the one behind Apple Intelligence as it ships in iOS 26 and macOS 26, is the one recent iPhones already carry. We ran all five of those on our own machines, so the charts group them as self-hosted, and every one of the 32 ran over the same 508 emails.

Before trusting any of these numbers, we checked the dataset itself, since it was not built for Pingbo. Its two annotators disagree most on its four-way labels, and the two-way question Pingbo is graded on lines up with who actually replied. On 24 of the 100 emails we had re-judged, a third reader disagreed with the dataset about it. The appendix walks through those checks for readers who want the method.

From a simple prompt to the full product harness

Pingbo is a model inside a harness we build ourselves, the form above, the checks on the answer, and the code that runs before any of it reaches you. A harness can make a model better or worse, and the only way to know which is to measure. So after every change we re-ran all 32 models over the same 508 emails, with nothing else changed, so any movement in the score came from our change and not from a different set of mail. Each round is a single run, and a model’s own answers shift a little between runs, so a difference of an email or two is not worth reading. Figure 2 is the whole record.

30%35%40%45%50%55%60%65%70%75%80%85%90%64.8% —Answering “no”to every emailAccuracy — share of the 508 emails sorted correctlyStep 0One of four labels,asked directlyStep 1The first prompts,measured end to endStep 2A verdict must quotethe message to claiman obligationStep 3The ask must usethe email’s ownwordsclaude-fable-5 — Step 0 82.5%, Step 1 66.3%, Step 2 83.9%, Step 3 84.3%; 508 emails in every arm82.5%66.3%83.9%84.3%claude-fable-5claude-fable-5-1 — Step 0 66.5%, Step 1 70.3%, Step 2 83.3%, Step 3 83.1%; 508 emails in every arm66.5%70.3%83.3%83.1%claude-fable-5-1claude-haiku-4-5-20251001 — Step 0 68.9%, Step 1 60.6%, Step 2 78.0%, Step 3 80.1%; 508 emails in every arm68.9%60.6%78.0%80.1%claude-haiku-4-5-20251001claude-opus-4-5-20251101 — Step 0 81.3%, Step 1 75.4%, Step 2 82.7%, Step 3 83.3%; 508 emails in every arm81.3%75.4%82.7%83.3%claude-opus-4-5-20251101claude-opus-4-6 — Step 0 77.2%, Step 1 67.7%, Step 2 80.5%, Step 3 82.7%; 508 emails in every arm77.2%67.7%80.5%82.7%claude-opus-4-6claude-opus-4-7 — Step 0 81.9%, Step 1 65.2%, Step 2 82.1%, Step 3 82.9%; 508 emails in every arm81.9%65.2%82.1%82.9%claude-opus-4-7claude-opus-4-8 — Step 0 81.5%, Step 1 72.8%, Step 2 84.8%, Step 3 84.6%; 508 emails in every arm81.5%72.8%84.8%84.6%claude-opus-4-8claude-opus-5 — Step 0 81.5%, Step 1 66.7%, Step 2 83.3%, Step 3 84.4%; 508 emails in every arm81.5%66.7%83.3%84.4%claude-opus-5claude-sonnet-4-5-20250929 — Step 0 67.9%, Step 1 68.7%, Step 2 81.3%, Step 3 81.7%; 508 emails in every arm67.9%68.7%81.3%81.7%claude-sonnet-4-5-20250929claude-sonnet-4-6 — Step 0 84.1%, Step 1 68.5%, Step 2 83.3%, Step 3 82.5%; 508 emails in every arm84.1%68.5%83.3%82.5%claude-sonnet-4-6claude-sonnet-5 — Step 0 78.7%, Step 1 71.7%, Step 2 83.9%, Step 3 83.7%; 508 emails in every arm78.7%71.7%83.9%83.7%claude-sonnet-5gemma-4-E4B-it-qat-4bit — Step 0 71.1%, Step 1 60.0%, Step 2 63.6%, Step 3 65.2%; 508 emails in every arm71.1%60.0%63.6%65.2%gemma-4-E4B-it-qat-4bitgpt-4.1 — Step 0 81.1%, Step 1 63.8%, Step 2 81.3%, Step 3 82.9%; 508 emails in every arm81.1%63.8%81.3%82.9%gpt-4.1gpt-4.1-mini — Step 0 71.1%, Step 1 54.5%, Step 2 72.8%, Step 3 75.6%; 508 emails in every arm71.1%54.5%72.8%75.6%gpt-4.1-minigpt-4o-mini — Step 0 78.0%, Step 1 34.4%, Step 2 46.5%, Step 3 75.8%; 508 emails in every arm78.0%34.4%46.5%75.8%gpt-4o-minigpt-5 — Step 0 76.8%, Step 1 60.4%, Step 2 78.9%, Step 3 80.5%; 508 emails in every arm76.8%60.4%78.9%80.5%gpt-5gpt-5-mini — Step 0 72.2%, Step 1 56.5%, Step 2 71.7%, Step 3 79.9%; 508 emails in every arm72.2%56.5%71.7%79.9%gpt-5-minigpt-5.1 — Step 0 79.5%, Step 1 59.4%, Step 2 82.1%, Step 3 82.5%; 508 emails in every arm79.5%59.4%82.1%82.5%gpt-5.1gpt-5.2 — Step 0 78.7%, Step 1 61.8%, Step 2 83.1%, Step 3 83.5%; 508 emails in every arm78.7%61.8%83.1%83.5%gpt-5.2gpt-5.4 — Step 0 82.5%, Step 1 67.9%, Step 2 82.5%, Step 3 82.9%; 508 emails in every arm82.5%67.9%82.5%82.9%gpt-5.4gpt-5.4-mini — Step 0 69.5%, Step 1 69.9%, Step 2 79.3%, Step 3 79.1%; 508 emails in every arm69.5%69.9%79.3%79.1%gpt-5.4-minigpt-5.5 — Step 0 82.5%, Step 1 72.2%, Step 2 83.9%, Step 3 85.0%; 508 emails in every arm82.5%72.2%83.9%85.0%gpt-5.5gpt-5.6-luna — Step 0 70.7%, Step 1 63.6%, Step 2 81.5%, Step 3 84.1%; 508 emails in every arm70.7%63.6%81.5%84.1%gpt-5.6-lunagpt-5.6-sol — Step 0 82.3%, Step 1 74.6%, Step 2 82.7%, Step 3 83.3%; 508 emails in every arm82.3%74.6%82.7%83.3%gpt-5.6-solgpt-5.6-terra — Step 0 79.9%, Step 1 67.1%, Step 2 82.9%, Step 3 83.7%; 508 emails in every arm79.9%67.1%82.9%83.7%gpt-5.6-terraornith-1.5-9b-mlx-full — Step 0 59.3%, Step 1 50.8%, Step 2 59.6%, Step 3 65.4%; 508 emails in every arm59.3%50.8%59.6%65.4%ornith-1.5-9b-mlx-fullqwen3.6-35b-a3b-mxfp4 — Step 0 75.0%, Step 1 65.0%, Step 2 73.4%, Step 3 79.7%; 508 emails in every arm75.0%65.0%73.4%79.7%qwen3.6-35b-a3b-mxfp4

OpenAI 16

  • gpt-5.6-luna
  • gpt-5.6-sol
  • gpt-5.6-terra
  • gpt-5.5
  • gpt-5.4
  • gpt-5.4-mini
  • gpt-5.4-nano
  • gpt-5.2
  • gpt-5.1
  • gpt-5
  • gpt-5-mini
  • gpt-5-nano
  • gpt-4.1
  • gpt-4.1-mini
  • gpt-4.1-nano
  • gpt-4o-mini

Anthropic 11

  • claude-opus-5
  • claude-opus-4-8
  • claude-opus-4-7
  • claude-opus-4-6
  • claude-opus-4-5-20251101
  • claude-fable-5-1
  • claude-fable-5
  • claude-sonnet-5
  • claude-sonnet-4-6
  • claude-sonnet-4-5-20250929
  • claude-haiku-4-5-20251001

Self-hosted 5

  • qwen3.6-35b-a3b-mxfp4
  • ornith-1.5-9b-mlx-full
  • ornith-1.5-9b-mlx-6bit
  • gemma-4-E4B-it-qat-4bit
  • apple-fm-on-device
Figure 2. Step 0 asks each model, with no pipeline around it, to pick one of the four labels the annotators used, and scores that answer on whether a reply is owed. The three rounds that follow are ours, measured on the same 508 emails each time, and every step ran across all 32 models we tested. The five struck through cannot beat the 64.8% you score by answering “no” to every email, so no line is drawn for them. On the model we run on our own hardware, making the system point to the email before it may claim you owe something was worth 8.5 points, and making it use the email’s own words a further 6.3. The dashed line is that same 64.8%. Hover, focus or click a line to follow one model.

Step 0 is the bare question, asked straight of each model with no harness at all. Each model reads one email and picks one of the four labels the annotators used, and we score the answer the way we score everything here, on whether it means a reply is owed. Half of the 32 models score above 77.0% that way, and 29 of them clear the floor, so the models can do most of this job unaided. It is not a product, though. A label puts no name on the chase, no draft under the card, and nothing on your board.

Step 1 is our first attempt at the whole form, and it is a fall. The moment the model had to fill in the form, it started inventing work. A colleague sends over a redraft and asks for nothing, and the system produces “reply with feedback or approval”. The median drops to 65.0%, which is the floor. Only 17 models still clear it, and against their own bare prompt our harness made 28 of the 32 worse. The models were not the problem. We were.

Step 2 is the fix for exactly that. To say you owe a reply, the system has to quote the words that asked for one, and code goes looking for that quote in the mail. No quote, no claim. The median rose to 81.3%, 24 models cleared the floor, and for the first time more of them did at least as well inside the harness as outside it, 23 of the 32.

Step 3 came from reading the quotes Step 2 was accepting. A fifth of them were not from the email at all. They came from the older messages quoted underneath it, the ones a reply carries along, so the system was treating a question the sender had been asked, and was now answering, as a question the sender was asking you. Others were closing pleasantries, or work handed over that belongs on your list but owes no message. So the ask now has to come from the sender’s own words in the message at hand. The median moved to 82.5%, 27 models clear the floor, the best of them scores 85.0%, and 25 of the 32 now do better inside the harness than outside it.

Step 3 is the round that matters most to you, and it is the least visible one in Figure 2. On the model we run on our own hardware it was worth 6.3 points, closing the gap to the rest of the field from 7.9 points to 2.8. So the quality of the call is Pingbo’s own, and it holds when the model underneath changes, rather than resting on whichever model is best this month. It was not free, though not in the way one number suggests. On that same model the stricter rule dropped the share of owed replies that got a reply card from 68.2% to 62.0%, and raised the share of reply cards that were really owed from 61.6% to 76.6%. Of the 19 owed replies that lost their reply card, 18 became To Act cards and stayed on the list, so what reached Waiting fell from 13 to 10. Every round is a trade like that, which is why all three numbers are measured again each time, so the trade is seen rather than guessed. And the harness does not help every model. Seven of the 32 still do better with the bare question, most of them the smallest or the most heavily compressed, which is why Pingbo’s model is chosen by weighing cost against measured performance rather than by price alone. Apple’s on-device model shows the limit most plainly. Asked the bare question it scores 70.1%, clear of the floor, yet inside the harness it falls below the floor at every step. At Step 3, 153 of the 508 emails do not fit in the 4,096 tokens, about 3,000 words, it can read at once alongside Pingbo’s instructions, and for 306 of the other 355 its first decision about where the mail belongs is one the rules reject. Not one email becomes a Matter, so it drafts no reply at all. A model has to be able to hold the harness before the harness can help it.

Step 3 is the pipeline Pingbo ships. What is left is not another step but a setting, and the setting is a choice about you.

Missing emails, or fewer interruptions

The same pipeline can be tuned quiet or thorough, and Figure 3 shows every model at each setting.

Shows all 50840030022020%40%60%80%100%50%60%70%80%90%100%Precision — how much of what it shows needed a replyRecall — of the 179 needing a replyclaude-fable-5 — To Respond 65/88%, +To Act 86/55%, +2nd read 92/56%65/88%86/55%92/56%claude-fable-5claude-fable-5-1 — To Respond 64/85%, +To Act 83/62%, +2nd read 88/61%64/85%83/62%88/61%claude-fable-5-1claude-haiku-4-5-20251001 — To Respond 59/83%, +To Act 83/53%, +2nd read 96/47%59/83%83/53%96/47%claude-haiku-4-5-20251001claude-opus-4-5-20251101 — To Respond 56/94%, +To Act 75/65%, +2nd read 88/61%56/94%75/65%88/61%claude-opus-4-5-20251101claude-opus-4-6 — To Respond 60/86%, +To Act 80/55%, +2nd read 92/53%60/86%80/55%92/53%claude-opus-4-6claude-opus-4-7 — To Respond 59/89%, +To Act 84/54%, +2nd read 93/53%59/89%84/54%93/53%claude-opus-4-7claude-opus-4-8 — To Respond 63/91%, +To Act 82/65%, +2nd read 90/62%63/91%82/65%90/62%claude-opus-4-8claude-opus-5 — To Respond 64/93%, +To Act 84/55%, +2nd read 92/55%64/93%84/55%92/55%claude-opus-5claude-sonnet-4-5-20250929 — To Respond 58/87%, +To Act 72/62%, +2nd read 92/53%58/87%72/62%92/53%claude-sonnet-4-5-20250929claude-sonnet-4-6 — To Respond 63/84%, +To Act 83/60%, +2nd read 90/58%63/84%83/60%90/58%claude-sonnet-4-6claude-sonnet-5 — To Respond 64/88%, +To Act 83/58%, +2nd read 93/55%64/88%83/58%93/55%claude-sonnet-5gemma-4-E4B-it-qat-4bit — To Respond 75/66%, +To Act 83/55%, +2nd read 96/50%75/66%83/55%96/50%gemma-4-E4B-it-qat-4bitgpt-4.1 — To Respond 63/86%, +To Act 87/57%, +2nd read 94/55%63/86%87/57%94/55%gpt-4.1gpt-4.1-mini — To Respond 77/63%, +To Act 89/47%, +2nd read 96/45%77/63%89/47%96/45%gpt-4.1-minigpt-4o-mini — To Respond 60/71%, +To Act 73/53%, +2nd read 93/52%60/71%73/53%93/52%gpt-4o-minigpt-5 — To Respond 70/74%, +To Act 86/51%, +2nd read 94/50%70/74%86/51%94/50%gpt-5gpt-5-mini — To Respond 70/72%, +To Act 85/45%, +2nd read 96/45%70/72%85/45%96/45%gpt-5-minigpt-5.1 — To Respond 59/87%, +To Act 86/54%, +2nd read 94/52%59/87%86/54%94/52%gpt-5.1gpt-5.2 — To Respond 65/85%, +To Act 87/55%, +2nd read 94/52%65/85%87/55%94/52%gpt-5.2gpt-5.4 — To Respond 56/93%, +To Act 77/64%, +2nd read 93/63%56/93%77/64%93/63%gpt-5.4gpt-5.4-mini — To Respond 54/82%, +To Act 83/56%, +2nd read 96/50%54/82%83/56%96/50%gpt-5.4-minigpt-5.5 — To Respond 66/89%, +To Act 84/62%, +2nd read 89/60%66/89%84/62%89/60%gpt-5.5gpt-5.6-luna — To Respond 64/89%, +To Act 84/56%, +2nd read 96/50%64/89%84/56%96/50%gpt-5.6-lunagpt-5.6-sol — To Respond 61/89%, +To Act 81/65%, +2nd read 88/64%61/89%81/65%88/64%gpt-5.6-solgpt-5.6-terra — To Respond 63/88%, +To Act 83/61%, +2nd read 94/59%63/88%83/61%94/59%gpt-5.6-terraornith-1.5-9b-mlx-full — To Respond 56/61%, +To Act 71/45%, +2nd read 98/45%56/61%71/45%98/45%ornith-1.5-9b-mlx-fullqwen3.6-35b-a3b-mxfp4 — To Respond 62/77%, +To Act 80/53%, +2nd read 92/51%62/77%80/53%92/51%qwen3.6-35b-a3b-mxfp4

OpenAI 16

  • gpt-5.6-luna
  • gpt-5.6-sol
  • gpt-5.6-terra
  • gpt-5.5
  • gpt-5.4
  • gpt-5.4-mini
  • gpt-5.4-nano
  • gpt-5.2
  • gpt-5.1
  • gpt-5
  • gpt-5-mini
  • gpt-5-nano
  • gpt-4.1
  • gpt-4.1-mini
  • gpt-4.1-nano
  • gpt-4o-mini

Anthropic 11

  • claude-opus-5
  • claude-opus-4-8
  • claude-opus-4-7
  • claude-opus-4-6
  • claude-opus-4-5-20251101
  • claude-fable-5-1
  • claude-fable-5
  • claude-sonnet-5
  • claude-sonnet-4-6
  • claude-sonnet-4-5-20250929
  • claude-haiku-4-5-20251001

Self-hosted 5

  • qwen3.6-35b-a3b-mxfp4
  • ornith-1.5-9b-mlx-full
  • ornith-1.5-9b-mlx-6bit
  • gemma-4-E4B-it-qat-4bit
  • apple-fm-on-device
Figure 3. All 32 models we measured, at each of the three ways Needs you can be built. The five struck through cannot beat the 64.8% you score by answering “no” to every email, so no curve is drawn for them. Up catches more of the email that needs you; left means more of what you are shown did not. The dashed diagonals mark equal amounts of mail shown, from 220 emails to all 508. Along one diagonal the amount shown is fixed, so a model with higher recall there also has higher precision and is judging better. A steeper diagonal means more shown, so the same recall comes with lower precision, or the same precision with higher recall. Hover or click a name or a line to follow one model.
What goes into Needs youCaught, of the 179 that owe a replyOf what it shows, owed a replyEmails shown
To Respond cards only119 (66%)89%133
Plus To Act cards — the setting Pingbo ships with150 (84%)62%241
Plus a second, independent read159 (89%)60%265

Table 1. Three ways to build the Needs you list, on the best model we measured, gpt-5.5, at optimisation Step 3. These are settings of one pipeline: every row is Step 3 configured differently. The last row catches the most, and shows just over half the inbox to do it.

The first row of Table 1 is To Respond cards only. The second adds the To Act cards, and it is the setting Pingbo ships with. The third adds a second, independent read of every Matter.

The two annotators recorded one thing, whether an email needed a reply. Needs you is wider than that, because it also picks up the emails that ask you to do something, and the annotators never marked those. The second row shows 241 emails where the first shows 133, so 108 more. The annotators agree that 31 of them owed a reply, which is why the catch column climbs from 119 to 150. The remaining 77 count as wrong, not because Pingbo was wrong to surface them, but because the annotators had no way to say “this one asks you to act”. Some of the 77 is exactly what Needs you is for and some is a real mistake, and nothing they recorded separates the two, so we claim neither. That is the whole of the fall from 89% to 62% in the middle column.

What Pingbo leaves off the list gets a number too. At the setting Pingbo ships with, the middle row of Table 1, gpt-5.5 shows 241 emails, so the other 267 stay out of Needs you. It caught 150 of the 179 that owed a reply, so 29 owed one and were left out anyway, and the other 238 were labelled as owing no reply. Bringing the 29 down is the work.

Those 267 are not one lane, and the labels do not say the emails needed nothing. Pingbo files 46 of them under Waiting and 124 under Done; 95 never became a Matter at all, landing in Reading or Screened; two are calls that failed. What the annotators recorded is whether a reply was owed, and nothing else, so no number here says whether one of those emails still wanted something of you.

On this model, at these settings, nothing wins on both counts at once: catching more means showing you more you did not need, and where to sit is a judgement, not a measurement. It is not a law. Five of the 32 gain on both counts when the second read is added, claude-opus-5 among them, because the read catches emails their pipeline had dropped without widening the list as much. The two mistakes are not equal. Miss an email that needed you, and somebody was waiting on an answer that never came, usually without you knowing. Show you one that did not need you, and you have lost the second it took to see that.

What Pingbo delivers

In this evaluation, we counted how many emails needing a reply Pingbo found and how many it left off Needs you. These 508 emails can help us find an approach that works for most people, but they cannot tell us what works for each user. For example, a startup founder may want to see every reply from an investor, however short. A company manager cc’d on emails all day may only want to see the few they need to handle personally. One standard cannot meet both needs.

The dataset behind every number here is imperfect too, as the appendix sets out. Even on the simpler question, whether an email needs a reply, its two annotators agreed only 70% of the time, so past a point a higher score on it would be measuring the labels rather than Pingbo. Real email is often more complicated, with longer threads, more people on them, and asks buried in forwards. What we are building toward is a Pingbo that is right for its owner 99.99% of the time. No matter how carefully we check datasets or how many new ones we collect, standard datasets alone cannot get us there.

Only Pingbo’s owner can, because only its owner knows what right means for their own mail. So Pingbo will draw the line for each owner and keep redrawing it. It will ask a few questions when you set it up, and then learn from what you do. A reply you send to something it left off Needs you is a signal it can read, wherever it had filed it, though not proof it was wrong, since a reply is not always owed. An email you move out of Needs you without replying is another, and no more certain, since you may have dealt with the thing it asked for. Read together and over time, signals like those move the line for you, so the trade in Table 1 stops being made once on our behalf and starts being made continuously on yours.

Pingbo is built as an agent that keeps adapting to its owner, and what matters most to you is the thing it is learning. When it says something needs you, it can show you who, what and why in the sender’s own words, and have the reply drafted in advance. When it tells you something can wait, you can set it aside for now without worrying. That way, you can spend your time and energy on what really needs your attention.

Appendix: How far to trust these numbers

The Sappelli dataset was not built for Pingbo. It sorts email four ways, while Pingbo decides whether an email belongs in Needs you or Waiting. Every number in this article is measured against this dataset, so before grading Pingbo on it we had to check two things, how far the four-way labels can be relied on and how the two-way question Pingbo is graded on compares. Two people agreeing is not the same as two people being right, and on the four-way question they had agreed on only 508 of the 1,145, 44% of the time. Where those labels were contested, a third reader mostly did not side with them. On the two-way question the same two people agree far more often, and their labels line up with who actually replied. Neither check says which answer is correct on any one email. The rest of this appendix is how we looked.

Early on the models looked wrong in a suspicious way. We ran all 32 over the same 508 emails, asking each to sort them the same four ways the annotators had. On average, any two of the eight that came closest to the dataset gave the same answer 78% of the time, and each of them matched the dataset’s answer only 63% of the time. A weak model gets things wrong in scattered ways, so it disagrees with other models about as much as it disagrees with the annotators. These models kept arriving together at an answer the annotators had called wrong. Models agreeing is not proof that they are right, since they can share a mistake, but it was enough to make those labels worth reading again.

So we had every disputed email judged again, blind. We took the 57 emails where all eight models contradicted the dataset, stripped out any clue as to which answer came from where, and mixed in 43 emails where the models and the dataset already agreed. The judge was itself a model, and one of the 32. The 43 show it read the sheets rather than answering at random. They cannot show it was not simply siding with machines, because on those 43 the models and the dataset say the same thing. Nothing was unsealed until every one had been scored. Figure A1 shows how the 100 came back.

The 57 emails they disagreed about4557If the judge had no real preference, the 50 it decided would look like this2525Judge went with the modelsJudge went with the datasetJudge went with neitherThe 43 they already agreed about — the check43With no real preference those 50 would split about 25 and 25. They came back 45 and 5.Flip a coin 50 times. A split of 45 to 5 or wider turns up about once in 240 million runs.It got all 43 right. Both sides agree there, so it cannot catch a bias to models.
Figure A1. The blind re-judging of the answers in dispute. Every sheet was decided without any clue which answer came from where, and nothing was unsealed until all 100 had been scored. The middle bar is what chance alone would look like on the 50 the judge decided. The judge was itself a model, and one of the 32 tested. The 43 controls show it read the sheets rather than answering at random; they cannot show it was not simply siding with the models.

On the disputed emails the judge sided with the models 45 times, with the dataset 5, and with neither 7. On the 43 where nothing was in dispute it agreed with the dataset every time. A split of 45 to 5 is as likely as flipping a coin 50 times and landing that lopsided or wider, about once in 240 million runs. Statisticians call that chance a p-value, and this one, from a binomial test, is 4.2e-9 against the usual bar of 0.05. Two cautions. Those 57 were the sharpest disagreements in the set, not a random sample, so this says where the two sides disagree most, not how often either is wrong elsewhere, and not which of them is right. And a sheet cuts a long email at 4,000 characters, which six of the 100 hit, so on those six the judge read less than the models did.

The dataset was not broken. It was being asked a four-way question its own annotators could barely agree on. Both of them labelled 1,115 of the 1,145 emails, and they disagreed on 607. In 138 of those, both said a reply was owed and differed only on whether it was owed now or later. In 134, both said none was owed and differed only on whether you still needed to read it. Those 272 arguments are about which kind, not about whether. On the two-way question, the same two people agree on 780 of the 1,115, 70% of the time.

Two people agreeing still does not make them right, so we checked the two-way labels against something neither annotator wrote, whether the recipient actually replied. The Enron archive keeps the mail that followed each email. A later message from one of its recipients, within 30 days, on the same subject or quoting the original, counts as a reply. That is a rough test. The archive carries no reply threading, so a recipient forwarding the mail elsewhere counts too, and 72 of the 115 replies we found rest on a matching subject alone. If a recipient’s own sent mail never made it into the archive, we cannot see whether they replied, and a missing reply proves nothing. So we first counted the 209 emails with a recipient on the archive’s own staff list.

If the labels had nothing to do with real replies, emails marked as needing one and emails marked as not would get replies about equally often, and the gap between them would be close to zero. It is not. Emails the dataset marks as needing a reply got one 51.8% of the time, against 24.8% for the rest, a gap of 26.9 points. That gap is estimated from a limited number of emails, so it comes with a 95% confidence interval, here 12.3 to 42.1 points, and even its low end is above zero. Counting every recipient who sends any mail in the archive, 491 emails in all, the gap narrows to 14.3 points, with an interval of 6.2 to 22.4, and its low end is still above zero. In other words, emails labelled as needing a reply really did get replies more often. That is an association across the set, not proof that any single label is right.

A third way of counting adds every email where a reply was seen, 257 in all, and widens the gap to 33.6 points. We report it last because that set is chosen partly by the answer being tested. The eight models’ majority answers showed smaller gaps on the same emails: 14.8 points on the 209, 23.0 on the 257, 5.5 on all 491. Two intervals that overlap do not settle which gap is the larger, so we compared the two email by email. The dataset’s gap is the larger one on all three sets, by 12.1, 10.6 and 8.8 points, and each of those differences carries an interval that stays above zero: 3.0 to 14.7 on all 491, and, on the two smaller sets, 0.6 to 24.2 and 0.3 to 21.2, which clear zero only just. An association is not the same as being right every time. A label records that the sender expected an answer, and a busy recipient may never send one, so an unanswered email does not prove its label wrong.

The two-way labels are not spotless either. On 24 of the 100 re-judged emails the judge’s own answer flips whether a reply is owed. A few emails appear more than once in the set, so those 24 cover 26 of the 508 rows we score, 5.1% of them. They are a count of the emails we found in doubt, not an estimate of how many labels are wrong: the 100 were chosen for disagreement, the judge is another model, and the other 395 emails were never read again. What they do say is that no score near the top of this dataset should be read as exact.

So, is this dataset worth measuring against? For its four-way labels, no, and nothing here grades on them: its own two annotators agreed on 44% of them, and where they disagreed most a third reader sided with the models 45 times to 5. For the two-way question every score in this article does use — does this email owe a reply — the checks came back the other way. The same two people agree on it 70% of the time. It lines up with what recipients actually did, on all three ways of counting who could be seen replying. And it lines up more closely than the eight models’ own majority answer does, which is the comparison that matters: labels no better than the models could not be used to grade them.

None of that makes any single label right, and no check here could: they are indirect, the 100 re-judged emails were chosen for being the most contested in the set, and 395 were never read again. What it does mean is that the dataset holds up well enough to measure with, at a resolution of a few emails rather than one — with 26 of the 508 rows in doubt, a model that leads another by an email or two does not lead it at all. That is the bar every number in this article should be read against.

Push things forward, quietly.

Pingbo is in beta, and we are choosing a first group of early users.

What you use today Pick any
How did you first hear about Pingbo? Optional · Pick one

If you’re chosen for an early group, we’ll email you how to install it, including the official Apple test build.