Correction, 2026-10-03: this post said the hypotheses came before any data and that the new classes went live with it. Neither was true. The tracker already held 134 Baiduspider, 42 Baiduspider-render and 4 DeepSeekBot requests, filed under other-bot, which I had not looked at; what they show is now in the section below, and two hypotheses are marked accordingly. The new classes reached the live tracker on 2026-10-03, when the older requests were also reclassified. H-CN2 also wrongly put GPTBot in the never-fires tier: it had fired the inline beacon seven times. And no Manus-User request has been seen; the sentence about it going into human described what would happen.

I nearly skipped this one. The obvious objection is audience: how many people in Europe, the US, Canada or Australia ask Kimi, Qwen or Doubao anything? Not many. If the question were “will Chinese AI assistants send you traffic”, the answer for most readers of this site would be “hardly any”, and that would be the end of it.

That is the demand side. On the supply side, Chinese companies crawl Western sites whether or not Westerners use their products. Alibaba, DeepSeek and Moonshot also release open-weight models (Qwen, DeepSeek, Kimi) that Western companies build on, so what those crawlers take from a JavaScript-heavy page ends up, indirectly, in products used here. Crawl volumes have been published before; Cloudflare has put out figures on Bytespider. I haven’t found anyone measuring whether these crawlers run JavaScript, fetch robots.txt, or see content that only exists after hydration. That part this site can measure.

So I’m running it as a side experiment with an end date, while the main work here stays on JS content patterns. Writing the plan down before looking at any data keeps me from quietly moving the goalposts later.

What the tracker has seen so far

Until this week, one China-based crawler had its own class: Bytespider, ByteDance’s training crawler. Its record here is thin. It made one request in the first week, for robots.txt only (OAI-SearchBot only fetched robots.txt). In 80 days it never fired the JavaScript beacon (JavaScript execution has two depths). Its name was also one of the costumes worn by a VPS address that claimed to be Google, OpenAI, Anthropic, Apple and Meta (the costume post).

Every other agent from China went into other-bot, the bucket for any user agent containing “bot”, “crawl”, “spider” or “fetch”. Manus-User would have gone into human; none has visited.

other-bot was not empty, though. By the time of this post it held 134 Baiduspider, 42 Baiduspider-render and 4 DeepSeekBot requests (tracker rows), unverified. Baiduspider-render fired the JavaScript beacon 42 times between 16 July and 19 September, on both channels, all from one Chinese network (AS4837). The 76 Baiduspider requests from Chinese networks (AS4837, AS23724) went straight to content and never asked for robots.txt; the other 58 came from VPS hosts used by the costume pool. All four DeepSeekBot requests came from one address in that pool, which on the same day also posed as GPTBot, ClaudeBot, Googlebot, Baiduspider and seven other crawlers.

19 new classes

Company Agent What it does Vendor documentation
ByteDance TikTokSpider training according to ai.robots.txt; other lists call it a link-preview fetcher none found
ByteDance DoubaoBot crawler for the Doubao assistant none, name only
DeepSeek DeepSeekBot training none found
Moonshot AI KimiBot, Kimi-SearchBot, Kimi-User training, search index, fetch on a user’s request yes, with an IP list per agent
Baidu Baiduspider, Baiduspider-render search index; the render variant runs JavaScript yes, verification by reverse DNS
Baidu ERNIEBot, YiyanBot training, fetch for the Yiyan assistant none, name only
Alibaba QwenBot, TongyiBot training, fetch for the Tongyi assistant none, name only
Alibaba YisouSpider Shenma search index none found
Huawei PetalBot Petal Search and Huawei’s assistant help page linked from the user agent
Huawei PanguBot training for the PanGu models none found
Zhipu ChatGLM-Spider unclear none found
Manus Manus-User browser agent acting for a user none I could read
Sogou (Tencent) Sogou spiders search index help page linked from the user agent
Qihoo 360 360Spider search index none found

Where the vendor describes an agent, the “What it does” column follows the vendor. Everywhere else it follows the ai.robots.txt list.

What the research turned up

Two of the ten companies document their agents in a way you can check. Moonshot has a crawler page listing the user agents of its three Kimi agents, with an IP list for each. Baidu documents Baiduspider’s user agents, the rendering one included, and says to verify them by reverse DNS. Huawei and Sogou put a help-page URL inside the user agent. For the other six I found nothing from the vendor, only third-party lists.

Five of the names have no real traffic behind them that I could find. DoubaoBot, ERNIEBot, YiyanBot, QwenBot and TongyiBot appear in blocklists and in publishers’ robots.txt files. I could not find a single real request carrying any of them, in public user-agent collections or in logs people have posted, and no vendor documents them. They may be real and rare. They may also be names copied from one blocklist to the next. They get classes anyway, so that if none of them turns up, the zero is a measurement.

Manus-User looks like a person. It sends a normal desktop Chrome user agent with ; Manus-User/1.0 on the end. There is no “bot”, “crawl”, “spider” or “fetch” in it, so an analytics filter built on those words counts it as a human visitor. Mine would have. Kimi-User has none of those words in its name either, but its user agent ends with a link to kimi.com/policies/kimi-crawlers, and “crawlers” is enough for a catch-all filter to catch it by accident.

Brand names catch browsers. A filter on “Sogou”, “Quark” or “baidu” also matches real people, because SogouMobileBrowser, the Quark browser and the Baidu app all carry those strings. The classifier matches each crawler’s own token instead (Baiduspider, 360Spider, Sogou ... Spider), and regression tests keep the Sogou, Quark and Baidu app user agents in human.

If you filter bot traffic today

  • Add Manus-User to your bot filter. Otherwise it counts as a human visit.
  • Match the crawler’s token. Baiduspider rather than baidu, Sogou web spider rather than Sogou.
  • Treat most of these user agents as claims. Kimi’s agents can be checked against Moonshot’s IP lists and Baiduspider by reverse DNS. For the rest, the user agent is all you have, and faking one costs nothing.

Five hypotheses, before the data

  • H-CN1. Most of these agents will not show up. next.jsseo.dev is an English site that went live in May 2026, and almost nothing in Chinese links to it. I expect at most a handful of the 19 classes to log a request in the first four weeks, and none of the five name-only agents.
  • H-CN2. The batch crawlers don’t run JavaScript. Bytespider, TikTokSpider, DeepSeekBot, KimiBot, PetalBot and the rest will land in the same “never fires the beacon” tier as ClaudeBot.
  • H-CN3. Baiduspider-render does. If it visits, the beacon fires. It is the only crawler on the list whose vendor says it renders pages. I am not predicting which of the two depths it reaches. (Already supported by the data above: it fired both.)
  • H-CN4. Kimi-User reads HTML only; Manus-User runs the page. Kimi-User will behave like ChatGPT-User, Claude-User and Perplexity-User did here: fetch the HTML and run nothing. Manus-User drives a real browser, so it should fire the beacon.
  • H-CN5. The documented crawlers fetch robots.txt before content. Kimi’s agents and Baiduspider will request /robots.txt before any content page. For the undocumented ones I have no prediction. (For Baiduspider, the unverified data above already runs against this.)

The plan

  1. Four weeks of passive observation. The new classes reached the live tracker on 2026-10-03, a week after this post, so the four weeks count from then. Nothing else changes: no Chinese content, no submission to Baidu’s or Petal’s webmaster tools. What arrives uninvited is the point.
  2. One afternoon of active probes. I’ll ask Kimi, Qwen and DeepSeek to read a fresh probe page built like the runtime-UUID probe that caught NotebookLM running our JavaScript. For whatever fetches it, if anything does, I’ll record the user agent, the network and whether the beacon fired. Doubao and Tencent’s Yuanbao may need a Chinese phone number to sign up; I’ll try anyway.
  3. Verification after that. None of the new classes is verified yet. Kimi’s IP lists and Baidu’s reverse DNS come next, once I have read both at the source. I have checked a crawler against the wrong list before (Anthropic publishes its crawler IPs) and would rather not repeat it.
  4. A decision at four weeks. If there is a signal, a findings post with the numbers. If there isn’t, a short follow-up that says so, and the classes stay in place.

What this will not tell you

  • One small English site. A crawler that ignores this site may be busy elsewhere. Volume here says nothing about volume on a large site or a Chinese one.
  • A user agent is a claim. Until verification is in, a “DeepSeekBot” request is a request that says it is DeepSeekBot. The costume post shows how cheap that claim is.
  • Only named agents get a class. A crawler that sends a plain browser user agent with no name in it lands in human here, as it would in most analytics.
  • A probe is one fetch on one day. The same product may do something else next week or for another user.

Method notes

The change is 19 patterns in the tracker’s classifier.js, placed before the catch-all. User-agent strings come from vendor documentation where it exists (Moonshot, Baidu) and otherwise from the public collections crawler-user-agents and Matomo device-detector; agent names and stated purposes also come from ai.robots.txt. Before shipping I ran the old and new classifier over about 5,400 real user agents from those collections. Every change moved a user agent from other-bot or human into its intended new class, and no existing class lost a single one. METHODOLOGY.md lists each class with its pattern and verification status.


Data availability: JS SEO Lab publishes methodology, tracker code, and classifier in the public repository at github.com/Qbeczek1/jsseo-dev. Live dashboard at /dashboard/.

Bias disclosure: I run JS SEO Lab as an independent technical SEO research project. I also do paid technical SEO and AI-visibility audits through FratreSEO. No framework vendor, crawler vendor, search engine, or AI company funds this work.