How to vet creators before outreach: the four-signal check that replaces your gut

By Yagna, founder of Sift · July 16, 2026

Here's the whole check. Everything after this box explains one line of it.

The four signals and the decision rule

  1. Campaign congruence. Does their content actually sit in your genre, or just adjacent to it? A fitness page isn't a recovery-tool audience.
  2. Creator credibility. Is the account alive, consistent, and growing the way a real creator grows — not the flat-then-spike shape of bought followers?
  3. Buyer-intent comments. Do the comments read like people who'd buy, or like a pod — arriving in a burst, generic, no reply threads?
  4. Product/demo content behavior. Not "have they shown this product" — they haven't, you haven't pitched them yet. It's whether their existing content style suggests they could carry a demo naturally if asked.

Accept requires 1, 2, and 3 all positive — 4 doesn't gate it. Reject is two or more of the four negative. Everything else is maybe/review.

An operator on Reddit put the problem better than any pitch deck: "even if you pay for fancy platforms you still end up using your gut to judge if someone is going to actually care about your product and deliver." Same thread, same person: landing 5-10 genuine partnerships meant reviewing twenty to thirty creators by hand, every campaign. That gut check — repeated thirty times per campaign, forever — is the job this post turns into a protocol.

Notice what's not in the box: follower count, engagement rate as a single number, verified badges. Those are the stats every discovery tool already hands you, and they're exactly the stats that lie. An account can clear every top-line metric on Modash and still fail congruence, credibility, or comments once you actually look. We didn't land on this shape on the first pass either — an earlier version rejected any creator who hadn't already posted a product demo, which cut good creators for the crime of not being asked yet. That fix is its own section, below.

The rest of this post walks each signal: what it catches, how to check it without a paid tool, and why one of the four works differently from the other three.

1. Why follower count and engagement rate stopped working

Comment-pod farms exist specifically to beat the metrics buyers check first. A pod is a group of accounts that comment on each other's posts within minutes of posting, every time, regardless of content. The effect: engagement rate looks healthy, comment count looks healthy, and none of it is a real audience reacting to anything.

So the tell isn't the rate — it's the shape. Two accounts can both show 4% engagement where one is a real audience and one is twelve handles rotating on a schedule, and no single number distinguishes them. Shape only shows up when you look at the comment section itself, sorted by time — which is exactly the step everyone skips, because it's the one step you can't automate by pulling an API field. Section 4 is the full checklist for reading it.

2. Signal 1 — Campaign congruence

Congruence fails quietly. A creator can be exactly the right genre and still be the wrong fit for this campaign, this product, this angle. The check is narrower than "do they post fitness content" — it's "would this specific audience buy this specific thing for the reason we're selling it."

This is the genre-mapping logic from the sourcing pipeline, run in reverse. Sourcing asks "what genres would use this product." Vetting asks, per candidate, "does this account's actual content match one of those genres, or does it just share a hashtag with one."

The Ultrahuman version makes it concrete. A hybrid-athlete account posting ice baths, HRV tracking, and sleep routines: pass — that audience buys a recovery wearable for exactly the reason we sell it. A bodybuilding page with twice the followers and the same #recovery hashtag: fail — same fitness shelf, but an audience watching for size and strength programming doesn't buy a sleep ring. Both clear every category filter on every discovery tool. Only reading the content catches the difference.

Evidence to pull: last 10-15 posts, read them, not the bio. The bio is marketing copy the creator wrote about themselves. The post history is what they actually make.

Run it today: take 10 candidates from your last sourcing batch, read each one's last 10 posts, and finish this sentence per candidate: "this audience buys our product because ___." Can't finish it — that's a congruence fail, whatever the hashtags say.

3. Signal 2 — Creator credibility

Credibility is about the account, not any single post. Two checks that catch what a snapshot metric can't:

Growth shape. Real accounts grow lumpy — flat stretches, then a spike when one video hits, then flat again. Bought-follower accounts show a suspiciously smooth ramp, or a single unexplained cliff-jump with no corresponding content spike. If the platform or tool shows a growth chart, this is a thirty-second look that catches what engagement rate can't.

Posting consistency and aliveness. Has this account posted in the last two to four weeks? Is there a pattern (weekly, twice a week) or does it look abandoned with one recent post to look active for a campaign push? A creator who posted daily for a year and went dark six weeks ago isn't dead by the numbers — they're dead by the check that matters, which is "will they actually deliver if we ship product."

Neither of these needs a paid audit tool. Both are visible on the profile in under a minute, once you know to look for shape instead of a single stat.

Run it today: for the same 10 candidates, check last-post date and eyeball the posting rhythm. Anyone silent past four weeks gets flagged now, before a package ships to a creator who already left.

4. Signal 3 — Buyer-intent comments

This is the direct answer to "how do I tell if the engagement is real." Read the comments, sorted by recency, on the last three to five posts. You're looking for:

  • Timing spread. Real comments arrive across hours or days. All-at-once within minutes of posting is the single strongest pod tell.
  • Reply depth. Real audiences argue, ask follow-up questions, get replied to by the creator or by each other. Pods don't reply to anyone, including the creator.
  • Specificity. "Does this actually work for oily skin?" is a buyer. "🔥🔥 love this" repeated by the same handles across five different posts is a pod.
  • Repeat handles across unrelated posts. Pull ten commenter handles from one post, check if the same set appears on a different creator's post in your candidate list. Pods rotate across accounts they're paid to inflate — if you're vetting fifteen candidates in the same niche, you'll catch the same names recurring.
  • Outlier check against the creator's own norm. Before trusting a comment section, check whether the post it's on is typical for this account or a one-off viral spike — compare that post's views/likes to the creator's usual range. A creator whose comments look great on their one viral post but thin everywhere else isn't reliably engaged; you're reading noise from an outlier, not their real audience.

That last check is the one that scales worst by hand and the one that matters most, because it's the only one of these that requires comparing across a creator's own post history, not just reading one comment section in isolation. It's also the reason we made Sift read actual comments instead of pulling an engagement-rate field: every verdict it returns cites the specific comments and posts it judged from, so you can spot-check the call instead of trusting a score.

Run it today: same 10 candidates, 3 posts each, comments sorted by time. Twenty minutes. If you've never done this deliberately, expect to catch your first pod before you're halfway through the list.

5. Signal 4 — Product/demo content behavior, and why it doesn't gate acceptance

This is the signal most people get wrong, and we got it wrong first. The instinct is to ask "has this creator ever shown a product on camera and talked about using it" — and reject if not. That's the wrong question, for an obvious reason once you say it out loud: you haven't reached out yet. Of course they haven't posted about your product. Grading a creator down for not having done the thing you're about to ask them to do is circular.

The right question is narrower and forward-looking: does this creator's existing content style — reviews, routines, before/afters, comparisons, storytelling, in any category — suggest they could carry a demo naturally if asked. That reframe still didn't fully fix it. Even asked that way, this signal kept coming back negative for creators who were, by every other measure, good fits — it has a structural bias toward "unproven" reading as "no," because most creators genuinely haven't demoed your specific thing yet, and a model (or a person, skimming fast) leans on the concrete evidence it has, which is the absence.

So the fix isn't a better prompt, it's a different place in the decision rule. Product/demo behavior stopped gating acceptance — a creator can be a clean accept with this signal unclear or even mildly negative, so long as congruence, credibility, and buyer-intent comments are all solid. It still counts toward rejection — if it's negative alongside one other negative signal, that's a real reject, because now you have two independent reasons to doubt the creator, not just an absence of proof on one axis. Asymmetric on purpose: unproven shouldn't block a good creator, but unproven-plus-another-red-flag should.

Run it today: pull your last batch of rejects and re-read the reasons. Every creator cut only for "no sponsored content yet" or "never demoed a product" goes back in the pile — that was the circular question, not a real signal.

6. Why "maybe/review" is a real bucket, not a cop-out

Most manual vetting collapses to binary: creator's fine, or creator's out. That's wrong for the same reason a wide sourcing net is right — a lot of real signal is genuinely ambiguous on first look, and forcing a binary call on ambiguous evidence either wastes good creators or lets bad ones through.

The rule: accept needs three signals clean — congruence, credibility, and buyer-intent comments all positive. Reject needs two or more of the four negative — a decisive double failure, not a single miss. Everything in between — one signal thin, one unclear, evidence too sparse to call — is maybe/review, and it stays in the pile for a second look or a lighter-touch outreach test, instead of getting thrown out because the first pass wasn't clean.

This matters more as volume goes up. At the review-30-to-land-10 rhythm from the operator quote up top, you can afford to eyeball everything twice. At the volume a real sourcing pipeline produces — hundreds of candidates per run — you need a bucket for "not clearly good, not clearly bad" that doesn't just default to reject, or you quietly filter out a chunk of your best candidates because the review was rushed.

7. The order that actually works

Run the math on doing this by hand. The four signals, checked honestly, take three to five minutes per creator. A single decent sourcing run returns 300+ candidates — that's 15 to 25 hours of vetting, a full work week, generated by one afternoon of sourcing. This is why "vetting is annoyingly manual" is the most common complaint from teams doing this at volume: the arithmetic is simply against them.

The wrong response to that arithmetic is to vet while you source, one candidate at a time, using vetting as the filter on what gets collected. Interleave them and every sourcing session turns into a vetting session — you source a tenth of what you could have, and you're still spending the same minutes per creator.

The shape that works: source wide and dirty first — the four-lane pipeline is built exactly to not filter as it runs — then vet in a separate, batched pass against the four signals above. Sourcing's job is finding candidates who are plausibly right. Vetting's job is deciding who's actually right. Different jobs, different rhythms.

This is the layer we built Sift for. It runs exactly the four signals in this post — congruence, credibility, buyer-intent comments, demo behavior — against your sourced list, at whatever volume the list is, and returns accept / maybe / reject per creator with the cited evidence behind each call. Not a black-box score: the verdict quotes the actual comments and posts it judged from, so the gut check you'd have done by hand is sitting there, already done, ready to be spot-checked.

The work week of manual vetting collapses back into the afternoon it should be. If you've got a creator list sitting in a sheet right now, start a trial and run it through — the first ranked shortlist tells you more than any pitch could.