Our crawler

This page is the address published in our crawler's user agent. It explains who operates it, what it does, and how to say no.

Current status: not crawling

Chuckles.today does not currently fetch anything from any external site. No web connector exists, no source has been approved for automated collection, and the ingestion kill switch in our database is off. The crawler has never run. If you see a request claiming to be ChucklesBot today, it is not us. This section will be updated before any crawling ever begins.

Operator and purpose

The crawler is operated by Chuckles.today, the curated humour publication you are reading. Its only purpose is discovering candidate comments from a short list of individually approved sources so that human editors can review them. It never publishes anything by itself: every published quotation passes context verification, a rights decision and human editorial approval first.

Its user agent is ChucklesBot/0.1 (+https://www.chuckles.today/crawler)

How it behaves

  • It always identifies itself. Every automated request we ever make will carry the user agent "ChucklesBot/0.1 (+https://www.chuckles.today/crawler)". We never impersonate a browser, a person or another crawler, and we never rotate identities or proxies.
  • It only visits approved sources. There is no general web crawler and never will be. Each source must pass a legal and editorial review (terms, robots, commercial-use permission, risk) and hold an approved, versioned collection policy before a single request is allowed. Unknown sites are simply never contacted.
  • It respects robots.txt and site terms. robots.txt is fetched and honoured before any request, and a disallow is treated as final. A source whose terms do not permit automated access is not collected from — we ask, licence, or drop the source.
  • It is rate-limited per source. Every source has its own reviewed request budget, recorded as requests per hour on its approved policy (never more than an average of one request per second). Requests to one source are made in sequence, never in parallel, and a 429 ends the pass.
  • It never bypasses access controls. No logins, no session or cookie reuse, no paywalls, no CAPTCHA solving, no private groups or members-only pages, no personal-user credentials. A source that requires authentication is blocked unless its owner has provided an official API or an explicitly authorised feed.
  • It stores the minimum and forgets on schedule. We store short excerpts and metadata, not page archives; full text only where a source's policy explicitly permits it. Every stored class has a declared retention window (raw payloads 14 days, excerpts 180, and so on), and content from an opted-out source is purged immediately with only a minimal record kept to prove the opt-out was honoured.

Opting out, complaints and takedowns

Any of these works, and none requires a reason:

  • robots.txt. Add the lines below to your robots file. They are honoured from the first request, and a source that blocks us is never worked around:
    User-agent: ChucklesBot
    Disallow: /
  • Ask us. A monitored contact address will be published on this page before the first automated request is ever made — publishing it is a precondition of switching the crawler on, not a follow-up. An opt-out disables the source, cancels queued work and purges unpublished content from it.
  • Individual authors. If you wrote a comment and want it removed from review or from a published edition, the same route applies without involving the site it appeared on, and takedown requests are handled by a human.

More on how the publication itself works is on the transparency page.