Learn Spam Protection Using Honeypot, Rate Limit and CAPTCHA

Oct 5, 2026SecurityEngineering
Summarize with AI: Google AI Claude ChatGPT Perplexity Grok

AI links open with a title + excerpt (these tools can't fetch the page themselves) — use "Copy full article" to paste the complete text for a fuller summary.

Share:

Introduction

This article demonstrates how we protect a public comment form, using four layers instead of one.

A comment form is the most exposed thing on a blog. It is anonymous, it writes to your database, and its whole purpose is to accept text from strangers. Every spam bot on the internet will find it.

Why Four Layers

No single check stops everything, and each one fails in a different way.

A CAPTCHA is strong, but it depends on a third party being reachable. A rate limit is cheap, but resets when your function scales. A honeypot catches simple bots for free, and clever ones walk straight past it. A moderation queue catches everything and costs you time.

Use them together and an attacker has to beat all four. More importantly, when one has a bad day the others still work.

Features

  • Free checks run first, expensive checks run last.
  • The database write happens only after everything passes.
  • Nothing spammy reaches a reader even if all the automated checks are beaten.

With the following steps, you can build the same thing.

  1. Add the honeypot field.
  2. Validate the input.
  3. Add the rate limit.
  4. Verify the CAPTCHA.
  5. Hold everything for moderation.

The order of those five is the design. Each step is more expensive than the one before it.

Add the Honeypot Field

A honeypot is a form field that is hidden from humans with CSS. A real visitor never sees it, so it is always empty. A bot fills in every field it finds.

if (body.website) {
  return { status: 201, jsonBody: { success: true } };
}

Look closely at what that returns. 201 Created. Success.

We did not save anything. But we told the bot we did.

This is the important part of a honeypot. If you return 400 Spam detected, the bot's author sees the failure, works out which field gave it away, and updates the bot to skip it. A fake success teaches them nothing. They move on believing the spam landed, and you never hear from them again.

It costs one hidden input and three lines, and it runs before anything else because it is the cheapest check you have.

Validate the Input

Normal field checks come next — still free, still no network, still no database.

if (!postSlug || !authorName || !authorEmail || !commentBody) {
  return { status: 400, jsonBody: { error: "'postSlug', 'authorName', 'authorEmail', and 'body' are required" } };
}
if (!EMAIL_PATTERN.test(authorEmail)) {
  return { status: 400, jsonBody: { error: "'authorEmail' is not a valid email address" } };
}

Everything is trimmed first, so a comment of three spaces counts as empty.

Add the Rate Limit

The rate limit keeps one source from flooding you.

const RATE_LIMIT_WINDOW_MS = 10 * 60 * 1000;
const RATE_LIMIT_MAX = 5;
const submissionsByIp = new Map<string, number[]>();

function isRateLimited(ip: string): boolean {
  const now = Date.now();
  const timestamps = (submissionsByIp.get(ip) ?? []).filter((t) => now - t < RATE_LIMIT_WINDOW_MS);
  timestamps.push(now);
  submissionsByIp.set(ip, timestamps);
  return timestamps.length > RATE_LIMIT_MAX;
}

Five comments per IP per ten minutes. The filter drops old timestamps on every call, so the map cleans itself and there is no timer to run.

Now the honest part. This is in-memory and per-instance.

Azure Functions can run several instances at once, and each one has its own Map. It also resets on cold start. So a determined attacker can get more than five through by spreading requests across instances, or by waiting for a scale event.

We know. It is a speed bump, not a wall. A real rate limit needs shared state — Redis, or a Cosmos container with a TTL — and that is a service to run, pay for and monitor, on a personal blog's comment form.

My suggestion: be clear in your own code about which of your defences are best-effort. Ours says so in a comment right above the constant. The danger is not having a weak layer, it is forgetting that it is weak and counting on it later.

Verify the CAPTCHA

Cloudflare Turnstile runs next, and it is fourth on purpose. It is the only check that makes an outbound HTTP call, so everything free runs before it.

const captchaOk = await verifyTurnstile(body.captchaToken ?? '', clientIp);
if (!captchaOk) {
  return { status: 400, jsonBody: { error: 'Captcha verification failed' } };
}

Note it is still before the Cosmos DB write. The database is the last thing to be touched by any request.

Hold Everything for Moderation

The last layer is the one that actually guarantees your readers never see spam.

const comment: Comment = {
  id: crypto.randomUUID(),
  docType: 'comment',
  postSlug,
  authorName,
  authorEmail,
  body: commentBody,
  status: 'pending',
  createdAt: new Date().toISOString(),
};

Every comment is saved as pending. The public endpoint only returns approved ones. So even a comment that beat the honeypot, the rate limit and the CAPTCHA still sits in a queue until a person approves it.

This is the layer that does not fail. The other three reduce how much lands in the queue; this one decides what reaches a reader.

The comment moderation queue in the admin area

Two details in the storage worth copying.

Comments live in the blog container, not their own, tagged with docType: 'comment'. That is a cost decision — the Cosmos account's shared throughput is already spent across the existing containers, and a fourth container would need its own minimum allocation. Blog queries filter on docType so the two never cross-match.

The email address is never returned publicly.

function toPublicComment({ authorEmail, docType, ...rest }: Comment): PublicComment {
  return rest;
}

A destructure that drops the two fields and returns the rest. The public type is built with Omit<Comment, 'authorEmail' | 'docType'>, so if somebody later returns a raw Comment from a public endpoint, TypeScript complains. The rule is enforced by the type, not by everyone remembering it.

The Order Is the Design

Put together, a spam request costs us almost nothing.

Layer Cost to us Stops
Honeypot 3 lines, no I/O Simple form-fillers
Validation A regex Malformed junk
Rate limit A map lookup Volume from one source
CAPTCHA One HTTPS call Most automation
Moderation Human time Everything else

The cheapest checks reject the most common attacks, so the expensive ones only ever see traffic that already looks real.

Conclusion

In this article we learned how to layer a honeypot, input validation, a per-IP rate limit, a CAPTCHA and a moderation queue on a public comment form, and why returning a fake success to a honeypot hit is better than an error.

If you only take one thing: order your checks by cost, and make sure the last layer is one that cannot be bypassed by any bot at all.

Reference

#spam#honeypot#rate-limiting#captcha#azure-functions#cosmos-db

Comments

Be the first to comment.

Leave a comment

Never shown publicly.

Comments are reviewed before appearing publicly.