Three agencies pitch you in the same week. All three show logos you recognise, a case study with an impressive meeting count, and a monthly number they sound confident about. Nothing in any of those decks is comparable to anything in the others, because none of them define their terms.
The takeaway up front: an outbound agency's promised output is a forecast, but its inputs are facts that exist before you sign — and the most reliable quality signal is how much of the input they will let you inspect while you can still walk away. A good agency shows you the list, names the sending domains, hands over a real sequence, and writes down what counts as a meeting. A weak one keeps all four vague and asks you to judge it on the number at the end.
Why the sales call tells you almost nothing
The pitch is the wrong evidence, and not because agencies are dishonest. The category has simply converged on three artefacts: familiar logos, a case study with a headline number, and a monthly meeting target.
The headline number is the emptiest of the three. Two agencies can both say they booked forty meetings in a quarter, where one means forty people who fit the client's profile, agreed to a specific time, and turned up, and the other means forty replies that sounded interested and were forwarded on. Neither is lying. Only one is selling pipeline.
The logos are barely better. A logo tells you a company paid an invoice. It does not tell you whether the engagement renewed, who ran it, or whether that person still works there.
So stop grading the promise. Grade the things that already exist.
Judge the four inputs, because they exist today
Everything an outbound programme produces comes out of four inputs. Each is inspectable before money changes hands, and each has a version that a weak agency works hard to keep abstract.
1. Where the list comes from
Ask how they would build your list, then ask to see a sample built to your criteria.
A good one describes its sourcing plainly, explains which filters it would use for your segment and why, produces a small sample against those filters, and asks for your exclusion list — current customers, open opportunities, accounts your team is already working — before it builds anything. It will also tell you which parts of your target market its data covers badly.
A weak one answers with volume. "We have access to two hundred million contacts" is a statement about a database, not about fit, and fit is the entire game. Watch for a refusal to produce a sample before contract, and for an inability to turn your ideal customer profile into concrete filters. If they cannot do that on a call, they will not do it in the account. Bring your own definition to the conversation — the method is in how to build an ideal customer profile.
2. Whose domain and inbox the mail leaves from
This is the input with a cost that outlives the contract, and the one buyers most often skip.
A good one registers dedicated sending domains for the campaign rather than pushing volume through your primary domain, explains its warm-up and volume ramp without being prompted, caps sends per inbox, sets authentication up properly, and will tell you exactly which domains it is using. It treats your main domain as something to protect, not a resource to spend.
A weak one says "we handle deliverability" and moves on. Two specific answers should end the conversation: refusing to name the sending domains, and proposing to send at volume from the domain your invoices and support email already run on.
The asymmetry is the point. If the engagement fails, the agency loses a client and you keep the reputation damage. Know what you are exposing before anyone touches your sending — the mechanics are in cold email deliverability.
3. Who writes the copy, and how much of it exists
Ask two questions: how many separate sequences will run for us, and who writes them.
A good one names the person, expects a small number of sequences aimed at distinct segments rather than one for everybody, shows you a real redacted sequence from another account, and wants your input on the offer. It will also tell you it plans to rewrite after the first few weeks, because the first version of anything is a hypothesis.
A weak one runs one template across every client with the industry noun swapped. You can usually detect this by asking what the last three sequences they wrote had in common. If the answer is a structure rather than a shared research method, you are buying a mail merge.
Ask, too, who replies. Copy quality is irrelevant if an interested reply sits unanswered for two days, and reply handling is where thin agencies quietly cut cost.
4. What counts as a meeting
This is the commercial term that decides whether the contract is worth anything, and it belongs in writing.
A good one already has a definition and will negotiate it: the role and company criteria the attendee must meet, agreement to a specific time, attendance, and a replacement policy for no-shows and for meetings that turn out to be a clear mis-fit. It is comfortable with the replacement clause because it does not expect to lean on it.
A weak one counts a positive reply, or counts a booked slot regardless of who is in it. Both produce a monthly report that looks healthy and a pipeline that does not move. Settle your own standard first — what "qualified" should mean before a lead reaches your calendar is a decision to make internally and then impose on any vendor.
What a good agency volunteers before you ask
Some signals only appear when nobody prompts them. Roughly in descending order of how much they tell you:
- It says no to something. A segment it does not think will work, a market it has no data for, a volume it will not send. Willingness to shrink the deal is the most expensive-to-fake signal there is.
- It sets a ramp expectation. New domains warm slowly, so the first weeks produce little. An agency promising meetings in week one is either planning to burn a domain or planning to redefine "meeting".
- It asks for your closed-lost list, not only your closed-won. The losses tell it who to avoid.
- It asks what happens after the meeting — who takes it, how often prospects fail to show, how fast you follow up. An agency indifferent to this is optimising for its own metric, not your revenue.
- The person pitching is the person running it, or introduces them before you sign.
What a weak one hides, and the phrases that hide it
- "Guaranteed meetings." Guarantees in this category are usually met by loosening the definition of the thing guaranteed. Ask what happens when the guarantee is missed, then read that clause twice.
- "Proprietary data." Sometimes real. Often a way to avoid saying where records came from. Ask how it is refreshed and how a removal request is handled; vagueness on either is your answer.
- A case study with no denominator. Meetings booked, but not from how many contacts, over how long, at what spend. A number with no denominator cannot be compared to anything, which is generally why it is presented that way.
- A minimum term longer than the ramp, with no exit. If the learning period sits inside a commitment you cannot leave, you cannot act on what you learn.
- No answer on ownership at the end. Who keeps the list, the reply history, and the domains when you part ways?
Structure the pilot so the answer shows itself
If the shortlist is genuinely close, stop analysing and design a trial that produces evidence instead of a report.
- Pick one narrow segment you can judge personally. You need to look at a booked meeting and know immediately whether it was a good one.
- Put the meeting definition in the agreement, with the replacement policy, before the first send.
- Ask for raw activity weekly — accounts touched, messages sent, replies verbatim — not a summary slide. The summary hides the pattern you need to see.
- Run it long enough to clear the warm-up. A pilot that ends before the ramp finishes has tested the ramp and nothing else.
- Agree the exit in advance: you keep the list, the reply history, and any domains registered on your behalf.
What stays yours regardless of who sends
An agency can run the mechanics. It cannot invent your ideal customer profile, fix a weak offer, or make your team turn up to the meetings it books. Those are the parts that decide whether outbound works at all, and no vendor is taking them off your hands.
Decide them first and vetting gets much easier, because you stop asking agencies what good looks like and start checking whether theirs matches yours. The underlying method is the same one you would run in-house, laid out in the sales prospecting playbook, and the same four inputs need the same scrutiny when you automate instead of outsource, covered in handing outbound to an agent.
FAQ
Should a small team outsource outbound at all?
Only once the offer is proven and someone internally has booked meetings with it. An agency amplifies a working motion; it cannot discover one for you. If nobody at your company has ever had a good conversation with a cold prospect, outsourcing buys volume against an untested message.
Is a performance-based deal safer than a retainer?
Not automatically. Paying per meeting shifts risk onto the agency, which is fair, but it also rewards booking anything that satisfies the letter of the definition. Performance pricing is only safe when the definition is tight and the replacement clause is real.
How long before an outbound engagement should be judged?
Long enough for domains to warm, sequences to run to completion, and at least one rewrite after the first real data arrives. Judging on the opening weeks measures setup. Judging on positive replies rather than qualified meetings measures nothing at all.
Should the agency send from my domain?
Generally no. Dedicated domains registered for the campaign keep cold volume away from the address your customers and invoices depend on. If an agency wants to send at volume from your primary domain, ask what its plan is for the reputation damage — and note that the plan would be your problem, not theirs.
Next step
You are not really evaluating agencies. You are evaluating four inputs — the list, the domains, the copy, and the definition of a meeting — and the good ones are simply those willing to show you all four while you can still say no. Write your own definition of a qualified meeting before the next call, decide which domains you will let anyone send from, and sort the shortlist against those two lines. Sharpen the underlying playbook at prospectuso.com.