New: the Unified InboxSee how
Back to Blog
Agency Operations • • 10 min read

How Many Mailboxes Per Inbox Manager? The Reply Desk Capacity Model for Agencies

How many mailboxes per inbox manager? There are three separate ceilings, each with its own math, and the one that binds first is almost never volume.

AW

Anirudh Walia

Founder & CEO

How Many Mailboxes Per Inbox Manager? The Reply Desk Capacity Model for Agencies

How Many Mailboxes Per Inbox Manager? The Reply Desk Capacity Model for Agencies

You just signed client twenty-six, your mailbox count crossed 1,500, and someone on the ops call asked the only question that matters: how many mailboxes per inbox manager can we actually run before replies start rotting in the queue? Nobody in the room has a number, so the decision gets made on vibes, which usually means you hire one person too late and find out from a client that a “yes, let’s talk” sat unanswered for two days.

That question has no single answer, and every post that gives you one flat ratio is guessing. What it has is three separate ceilings. Each one has real math behind it, each one caps your desk independently, and the one that binds first for most agencies is not the one everybody plans for.

Mailboxes are not the unit of work

Mailboxes do not create work. Replies create work. A mailbox sending 25 emails a day into a well-targeted list generates a different workload than one sending 25 into a scraped list of the wrong titles, and the headcount you need tracks the second number, not the first.

So before you can answer the ratio question you have to convert mailboxes into minutes. Four inputs, all of which you already have in your sequencer:

  1. Mailboxes under management (M). Count the live ones, not the ones still warming.
  2. Sends per mailbox per day (s). On 2026 deliverability settings most agencies sit between 20 and 40.
  3. Total reply rate (r). Every reply, not just the good ones. Out of office, autoresponders, “remove me”, and “who is this” all land in the queue and all cost someone attention.
  4. Reply mix. What share of those replies needs a human-quality written answer versus a two-second classification.

Daily replies are M × s × r. That is the top of your funnel of work. Everything else is how you spend it.

The three ceilings that cap mailboxes per inbox manager

Ceiling 1: volume, measured in minutes of handle time

This is the one everyone models, and it is the easiest. Assign a handle time to each reply type and add it up.

Realistic handle times for a competent inbox manager working inside a unified inbox, reasoned from what the task actually involves rather than from a stopwatch:

  • Clear positive, needs a booking push: 4 to 6 minutes. Read the thread, confirm which client and which offer, check the persona’s calendar rules, write in that sender’s voice, send, log it.
  • Soft or neutral reply (“send info”, “not right now”, “what is this”): 2 to 4 minutes. Shorter writing, same context lookup cost.
  • Negotiation, objection, or multi-stakeholder thread: 8 to 15 minutes, and sometimes a Slack message to the client first.
  • Everything else: 5 to 15 seconds of triage, times a very large number.

Then divide by productive capacity, not shift length. An inbox manager on an eight-hour shift does not get eight hours of reply handling. Between standups, client Slack, CRM hygiene, lead list cleanup, and the cost of switching between four clients with four different offers, 5 to 6 hours of genuine handle time is a good day. Model 5.5.

Ceiling 2: context, measured in clients per brain

This is the ceiling that actually binds, and almost nobody staffs for it.

To answer a reply well, an inbox manager has to hold a surprising amount in their head: the client’s offer and how it is positioned, the ICP and which titles are worth the calendar slot, the approved answer to the pricing question, the competitors they are allowed to name, the sender persona’s voice, the booking rules, and the specific thing this client got annoyed about last month. That is not a document you look up. It is working memory, and it is per client.

Past roughly five to eight concurrent clients, that memory degrades, and it degrades silently. The manager does not slow down. They start writing replies that are technically fine and subtly wrong: the generic answer instead of the client’s answer, the wrong confidence level on a pricing question, a meeting booked with a title that client does not sell to. Nobody catches it because the queue looks healthy.

The context ceiling is a count of clients, and it is largely independent of volume. Six clients at 25 mailboxes each is 150 mailboxes and a full context load. Four clients at 90 mailboxes each is 360 mailboxes and a lighter one. Same person, same brain, 2.4x difference in the ratio.

Ceiling 3: SLA, measured in peak concurrency

Here is where average-load staffing quietly fails. Replies do not arrive evenly. They cluster hard in the prospect’s morning, and if you are selling across US and EU timezones you get two clustered peaks inside a long coverage window.

Suppose 45% of a day’s actionable replies land inside a three-hour window. Staffing to the daily average means that window is structurally underwater, and the queue you built up at 9am does not drain until 2pm. An average handle time that looks fine on a spreadsheet produces a four-hour wait on the reply that mattered most, because the good replies arrive in the same crush as everything else.

This matters because speed to first response is the single highest-leverage variable in reply-to-meeting conversion. Intent decays fast, and a positive reply answered at hour four is competing against whoever answered at minute five.

And a tight SLA is not a throughput problem, it is a concurrency problem. To answer within five minutes you need a human who is free at the moment the reply lands. That means staffing above the mean, which means a desk that is idle most of the day and still occasionally late. Queueing behaves this way regardless of how good your people are.

A worked example: 1,500 mailboxes across 25 clients

Run the model with concrete numbers. Substitute your own, the shape is what matters.

Inputs: 25 clients, 60 live mailboxes each, 1,500 mailboxes total. 25 sends per mailbox per day, so 37,500 sends per day. Total reply rate 3%, so 1,125 replies per day.

Reply mix (typical shape for cold outbound, and the first thing you should measure against your own data):

Reply typeShareCount per dayHandle timeMinutes
Auto-replies, OOO, routing noise35%39410 sec66
Hard negatives and opt-outs30%33810 sec56
Soft and neutral, needs a written answer25%2813 min843
Clear positive interest10%1125 min560
Total100%1,1251,525

Ceiling 1, volume. 1,525 minutes of handle time per day, divided by 330 productive minutes per manager, is 4.6 FTE. That puts the volume ceiling at roughly 325 mailboxes per inbox manager.

Ceiling 2, context. 25 clients across 4.6 managers is 5.4 clients each, which lands right at the edge of comfortable. This configuration happens to be balanced. Now change one input: if those same 1,500 mailboxes were spread across 50 clients at 30 mailboxes each, volume math still says 4.6 people, but context math says 50 clients divided by a cap of 6 is 8.3 people. The volume ceiling becomes irrelevant. You are hiring for client count, not for reply count, and if you budget on the volume number you will be short by four people and never understand why quality is slipping.

Ceiling 3, SLA. Of the 393 replies that need a written answer, 45% arriving in a three-hour window is 177 replies carrying roughly 632 minutes of work into a 180-minute window. That is 3.5 people doing nothing but replying for those three hours, just to keep the queue flat, and keeping a queue flat is not the same as hitting a five-minute SLA. Stretch coverage across US and EU business hours and the same 4.6 FTE of work now has to be spread across a 14-hour window, which is a scheduling problem, not a headcount problem, and it is how agencies end up with one person covering 1,500 mailboxes at 7am.

The binding ceiling here is SLA, and the answer to “do we need another inbox manager” is yes, but adding one will not fix the thing the client is complaining about. Another body raises average throughput. It does not put someone free at the keyboard the instant a positive reply lands.

So what is the number?

Taking the model and varying the inputs that actually move it:

Desk configurationSLA on positivesMailboxes per inbox manager
3 to 4 clients, one offer shapeNext business day400 to 600
3 to 4 clients, one offer shape1 hour250 to 350
8+ clients, distinct offers and ICPs1 hour100 to 180
15+ clients, distinct offers and ICPs1 hour60 to 120
Any configurationUnder 5 minutesNot reachable with humans alone
Agent-first desk, humans review exceptionsUnder 5 minutes1,000+

Two things to take from that table. First, the spread is roughly 10x, which is why any single published ratio is useless. Second, the row that breaks the pattern is the five-minute SLA, and it does not break because of effort or talent. It breaks because arrival clustering and a five-minute promise are mathematically incompatible with a staffing model that bills by the hour.

How an AI inbox agent moves every ceiling

This is the part of the model most agencies get backwards. They treat an AI inbox agent as a productivity tool that makes each manager a bit faster, so they expect the ratio to move 20 or 30%. It does not work that way. The agent does not raise the ceilings. It removes two of them and turns the third into a configuration value.

Volume ceiling. The 732 noise replies per day never reach a human at all. Classification happens on arrival, opt-outs are processed, OOO replies are rescheduled against the return date, and nothing lands in a person’s queue unless a person is required. The 281 soft and neutral replies are the ones an agent handles best, because they are high volume and low variance: the right answer to “send me some info” is a known answer per client, and writing it 281 times a day is exactly the work humans do worst.

Context ceiling. This is the big one. The context ceiling exists because a human cannot hold 25 clients’ offers, ICPs, guardrails, and voices in working memory. An agent configured per client holds each one without degradation, and holds it identically at 9am and at 6pm on a Friday. The human’s context load collapses from “every client’s full operating manual” to “the exception in front of me, with the client context attached to it.” Clients per brain stops being the constraint because the brain is no longer the thing holding the clients.

SLA ceiling. Concurrency becomes free. 177 replies arriving in a three-hour crush is a staffing emergency for a human desk and a non-event for an agent, because there is no queue. Every one of them gets a first response inside five minutes, including the one at 7am and the one at 11pm. The five-minute SLA stops being a promise you staff for and becomes a setting.

Re-run the worked example on an agent-first desk. The human-touch set is no longer 393 replies, it is the escalation set: the negotiations, the multi-stakeholder threads, the genuinely unusual, and whatever your escalation rules flag for review. At 15% of the previous human set, that is roughly 59 conversations a day at 5 minutes each, which is 295 minutes, which is one person covering all 1,500 mailboxes with room left over. The ratio moves from 325 to 1,500, and it moves because the job changed from answering replies to reviewing exceptions and closing the hard ones.

That is also why the role itself improves. An inbox manager spending their day on 59 real conversations instead of 1,125 queue items is doing work that is actually worth $4,000 a month, which is a better business for you and a better job for them.

How to measure your own ratio in two weeks

Do not adopt the numbers above. Measure four things and run your own model:

  1. Total daily replies per 100 mailboxes. Pull it from your sequencer. This single number tells you more about your targeting than your reply rate does.
  2. Your actual reply mix. Sample 200 replies and classify them by hand into noise, negative, soft, and positive. Most agencies are surprised by how large the noise bucket is, and it is the cheapest thing to automate away. If you have not built this taxonomy yet, start with a reply triage workflow.
  3. Median and 90th percentile time to first response, split by reply type. The median will look acceptable. The 90th percentile on positive replies is the number that is costing you meetings, and it is the number your client will eventually quote back to you.
  4. Arrival distribution by hour. Bucket a week of replies by hour of arrival. Your peak-to-average ratio is your real staffing multiplier, and it is the input nobody collects.

With those four numbers you can answer the headcount question in a spreadsheet instead of on a vibe, and you can tell which of the three ceilings you are actually hitting. If it is context, another hire helps. If it is SLA, another hire mostly does not.

Hire, outsource, or deploy an agent

A straightforward way to decide, once you know your binding ceiling:

  • Volume ceiling binding, few clients, relaxed SLA. Hire. The work is linear and a person absorbs it fine. Check the cost per reply math before you commit to a salary, but headcount is a legitimate answer here.
  • Context ceiling binding, many clients with distinct offers. An agent first, then hire. Adding humans to a context problem multiplies the number of brains that each hold an incomplete picture. Standardize per-client context in a system before you add people to it.
  • SLA ceiling binding. An agent, and it is not close. No staffing model delivers a sub-five-minute first response across a 14-hour coverage window at a cost that survives contact with your margins.
  • Growing faster than you can hire, or you do not want to run a reply desk at all. Underfive runs a fully managed inbox service for agencies, where the agent handles the volume and our team owns the exceptions across your client book.

The honest summary of “how many mailboxes per inbox manager” is this: with a human-only desk, somewhere between 60 and 600 depending on client count and SLA, and you should calculate it rather than guess. With an agent handling first response and triage, the ratio stops being a staffing constraint at all, which is the whole point. The reply desk is the one function in outbound where throughput and speed are the same problem, and headcount only ever solves one of them.

Sequencers send. The conversation layer is what turns those replies into booked meetings, and it does not get tired at 6pm on a Friday. Run your four numbers, find your binding ceiling, and if it is SLA or context, see what the agent-first model does to your ratio.

mailboxes per inbox manager cold email inbox manager reply desk staffing inbox management at scale agency operations AI inbox agent agency reply desk

Share this article

AW

Written by

Anirudh Walia

Founder & CEO

Ready to reply faster?

Underfive responds to your leads in under 5 minutes, 24/7. Start converting more leads today.

Book a Demo