Bytespider is one of the crawlers site owners ask about most, and one of the least documented. ByteDance publishes no page we could reach that describes it, so everything below comes from independent reports, each named with its date.
What is known about Bytespider
- Operator. The user agent strings that others record carry ByteDance and Toutiao addresses. ByteDance is TikTok's parent company.
- Purpose. Not documented by ByteDance on any page we could reach. Third-party pages claim search and AI training uses without citing ByteDance.
- robots.txt name. Everyone uses
Bytespider. Whether the crawler honors it is disputed, as below. - IP list. None from ByteDance that we could reach.
Does Bytespider follow robots.txt?
Several independent reports say it doesn't, at least not reliably:
- Fortune, October 2024. Fortune reported research showing that Bytespider didn't respect robots.txt, and quoted Kasada's CEO saying it had been scraping data at about 25 times the rate of GPTBot. The same article made a similar claim about OpenAI's and Anthropic's crawlers, whose own documentation says GPTBot and ClaudeBot follow robots.txt.
- HAProxy, October 2024. HAProxy wrote that some AI crawlers, Bytespider included, don't identify themselves transparently, try to pretend to be real users and ignore robots.txt. Close to 90% of the AI crawler traffic HAProxy saw came from Bytespider.
- TollBit, first half of 2026, as reported by Search Engine Journal in August 2026: ChatGPT-User, Bytespider and YouBot each accessed disallowed pages on nearly half of the European sites that had explicitly listed them.
Its volume has dropped since 2024. Cloudflare reported in July 2024 that Bytespider led the AI bots it saw in requests, in how much of the web it crawled and in how often it was blocked. In July 2025, Cloudflare reported that Bytespider's request volume had fallen 85%, from second to eighth place in crawler share, at 2.9%.
How to block Bytespider
Start with robots.txt, which costs nothing:
User-agent: Bytespider
Disallow: /
Because the reports above say robots.txt isn't reliably honored, back it with a firewall or WAF rule that blocks requests whose user agent contains Bytespider. The rule doesn't depend on the crawler's cooperation, and it also stops impostors that use the name, which does no harm here.
Blocking Bytespider doesn't affect any ByteDance service we know of, since ByteDance documents none that depends on it.
User agent strings
ByteDance publishes no strings that we could reach. Third parties record different ones, for example:
Mozilla/5.0 (compatible; Bytespider; [email protected])
Mozilla/5.0 (compatible; Bytespider; https://zhanzhang.toutiao.com/) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.0.0 Safari/537.36
Match on Bytespider rather than on the whole string.
Can you verify Bytespider?
No. Without an IP list or a reverse DNS method from ByteDance, a request that says Bytespider can't be confirmed as genuine. The third-party IP lists that circulate are old or unverified.
Where to see Bytespider
Google Analytics 4 automatically excludes traffic from known bots and spiders, and the exclusion can't be turned off, so Bytespider won't appear there. Search your server and CDN logs for Bytespider.
OneLence AI visibility recognizes Bytespider by its user agent when your server reports page requests to OneLence, and shows which pages it requested and how often, marked as matched by user agent because ByteDance publishes no IP list.
Frequently asked questions
What is Bytespider?
Bytespider is a web crawler associated with ByteDance, TikTok's parent company: the user agent strings people record carry ByteDance and Toutiao addresses. ByteDance publishes no documentation we could reach on what Bytespider collects or why.
Does Bytespider respect robots.txt?
Not reliably, according to several independent reports. Fortune and HAProxy reported in 2024 that it ignored robots.txt, and TollBit's report on the first half of 2026 found Bytespider accessing disallowed pages on nearly half of the European sites that had explicitly listed it. ByteDance publishes no statement we could check.
How do I block Bytespider?
Add a robots.txt group with User-agent: Bytespider and Disallow: /, and back it with a firewall or WAF rule that blocks requests whose user agent contains Bytespider. The firewall rule doesn't depend on the crawler's cooperation, and it also stops impostors that use the name.
Is Bytespider used to train AI?
ByteDance doesn't say on any page we could reach. Third-party pages claim search and AI training uses, but they cite no ByteDance statement, so treat those claims as unconfirmed.
Can I verify that a request really came from Bytespider?
No. ByteDance publishes no IP list and no reverse DNS method that we could reach, and the third-party IP lists that circulate are old or unverified. A request that says Bytespider could come from anyone.
