Công cụ miễn phí · không cần tài khoản

Free Robots.txt Tester & Validator

Test and validate robots.txt rules, User-Agent directives, blocked paths and AI bot access policies for search engines and AI crawlers.

Còn được gọi là: robots.txt tester · validator
Bằng chứng được hiển thị với nguồnKhông cần tài khoảnKhông có chỉ số nào được tạo ra.

Chạy một kiểm tra thực tế ở trên

Gửi một URL công khai, tên miền, hoặc từ khóa. Novaverb sẽ chỉ hiển thị bằng chứng mà công cụ này thực sự có thể lấy hoặc đo lường.

Mô hình bằng chứng

Biết điều gì mà kết quả chứng minh

This check proves what a site's robots.txt actually says and which rules apply to which crawler, including conflicting or overly broad Disallow blocks and any declared Sitemap directives. It proves crawl permission, not indexing: a URL that is disallowed here can still appear in search results, and one that is allowed is not thereby crawled.

1. Nguồn

Live robots.txt fetch. The result identifies where its evidence came from.

2. Ranh giới

The checker fetches /robots.txt and reads its directives. robots.txt is a crawl directive, not access control or index removal; this tool does not simulate a specific crawler's policy.

3. Hành động tiếp theo

Sử dụng phát hiện để xác minh một vấn đề, sau đó kết nối một không gian làm việc khi bạn cần lịch sử, giám sát hoặc phân tích toàn bộ trang web.

Giải thích Robots.txt

Kiểm tra robots.txt của bạn mà không tuyên bố quá mức những gì nó làm

robots.txt cho các trình thu thập tuân thủ biết các đường dẫn mà họ có thể yêu cầu. Đây là tệp đầu tiên mà hầu hết các trình thu thập lấy, vì vậy một dòng sai có thể giữ toàn bộ trang web ra khỏi tìm kiếm - hoặc cho phép các đường dẫn riêng tư được thu thập.

Điều gì cần kiểm tra

  • Địa điểm - tệp phải sống ở gốc miền và trả về HTTP 200 dưới dạng văn bản thuần túy.
  • Nhóm tác nhân người dùng - các quy tắc Allow/Disallow của mỗi nhóm chỉ áp dụng cho các tác nhân mà nó chỉ định; một Disallow trống sau User-agent: * có thể chặn mọi thứ.
  • Sơ đồ trang đã khai báo - xuất bản URL tuyệt đối của sơ đồ trang tại đây để các trình thu thập có thể phát hiện nó.
  • Ranh giới xác thực - trang này kiểm tra khả năng tiếp cận HTTP và bằng chứng chỉ thị; nó không phải là một trình phân tích RFC 9309 đầy đủ hoặc một kiểm tra Search Console.

Tại sao điều này quan trọng cho SEO và phát hiện AI

Một robots.txt rõ ràng giảm thiểu sự mơ hồ cho các trình thu thập thông tin tìm kiếm và AI. Novaverb có thể sử dụng tín hiệu này như một phần của mô hình thu thập và bằng chứng kỹ thuật rộng hơn, bên cạnh khả năng lập chỉ mục, canonicals và liên kết nội bộ.

Nhận câu trả lời đúng

Điều gì cần gửi - và điều gì cần tránh

Submit the domain, or any URL on it; the host is resolved and robots.txt is fetched from the root, which is the only location the specification recognises. A file placed in a subdirectory is not a robots.txt and will be reported as missing, which is the same conclusion every real crawler reaches.

Sử dụng nó như thế này
yourdomain.comBất kỳ hình thức địa chỉ nào cũng hoạt động - chúng tôi lấy /robots.txt tại gốc máy chủ cho bạn. http hoặc https, có hoặc không có www, một miền trống hoặc một đường dẫn đầy đủ - chúng tôi chuẩn hóa cho bạn.
https://www.yourdomain.com/anythingNgay cả một URL sâu cũng tốt; chúng tôi giải quyết đến máy chủ và đọc robots.txt gốc của nó.
Tránh điều này
Expecting a subfolder robots.txt to countrobots.txt chỉ được tôn trọng tại gốc máy chủ; một bản sao bên trong một thư mục con sẽ bị các trình thu thập thông tin bỏ qua.
Phương pháp công khai

Chính xác cách kết quả này được sản xuất

The host is resolved from the input and robots.txt is fetched from its root. User-agent, Allow and Disallow groups are parsed per RFC 9309, and the rules that apply to each major crawler are resolved the way that crawler resolves them. Conflicting or overly broad Disallow blocks are flagged, and Sitemap directives are noted.

  1. Chúng tôi giải quyết máy chủ từ đầu vào của bạn và lấy robots.txt từ gốc của nó.
  2. Chúng tôi phân tích User-agent, Allow và Disallow theo RFC 9309 và xác định các quy tắc áp dụng cho các trình thu thập chính.
  3. Chúng tôi đánh dấu các khối Disallow mâu thuẫn hoặc quá rộng và ghi chú bất kỳ chỉ thị Sitemap nào.
Xây dựng trên các tiêu chuẩn công khai

Các tiêu chuẩn quốc tế mà kiểm tra này áp dụng

RFC 9309 has defined the Robots Exclusion Protocol as a published standard since 2022, and this parse follows it rather than folklore, including its precedence rules for Allow against Disallow. Google's own robots.txt documentation is applied where it defines behaviour the RFC leaves to the implementation.

IETFRFC 9309
Robots Exclusion Protocol

Phân tích các chỉ thị Allow / Disallow / user-agent của bạn theo tiêu chuẩn robots.txt chính thức.

Đọc thông số kỹ thuật
Googlerobots.txt
Google robots.txt specification

Ghi chú nơi mà sự diễn giải của Googlebot mở rộng tiêu chuẩn cơ bản.

Đọc thông số kỹ thuật
Chúng tôi chỉ liệt kê một tiêu chuẩn khi công cụ này thực sự đọc hoặc đo lường theo đó. Khi một tín hiệu nằm ngoài kiểm tra trực tiếp, kết quả sẽ nói rõ điều đó thay vì ngụ ý rằng có sự bao phủ.
Câu hỏi thường gặp

Robots.txt Checker Câu hỏi thường gặp

What does the Robots.txt Checker check?

It fetches your /robots.txt file, reports the HTTP status, and lists crawl rules grouped by user-agent, counting every Allow and Disallow directive plus any declared Sitemap lines so you see exactly what crawlers are told.

What is robots.txt?

Robots.txt is a plain-text file at your domain root that follows the Robots Exclusion Protocol (RFC 9309). It tells cooperating crawlers like Googlebot which paths they may or may not request, grouped by user-agent.

Why does robots.txt matter for SEO?

Robots.txt controls crawler access to your paths. A mistaken Disallow can stop search engines from crawling important pages, so verifying the rules prevents accidentally hiding content you actually want discovered and ranked.

Does robots.txt remove a page from Google?

No. Robots.txt only requests that crawlers skip a path; it is a crawl directive, not index removal. A disallowed URL can still be indexed from external links. Use a noindex meta tag or removal request instead.

How do I stop robots.txt from blocking my whole site?

Look for a Disallow: / line under User-agent: *, which blocks every path. Remove or narrow it, then re-check. The tool flags this specific site-wide block so you can catch it fast.

What is a good HTTP status for robots.txt?

A 200 OK means the file was served and its rules apply. A 404 means no file exists, so crawlers assume full access. Persistent 5xx errors can cause crawlers to pause crawling entirely.

What's the difference between Allow and Disallow?

Disallow lists paths crawlers should not request; Allow re-permits a sub-path inside a broader Disallow. The most specific matching rule wins, so Allow can carve exceptions out of a blocked directory.

Should robots.txt list my sitemap?

Yes. A Sitemap: directive with the full URL helps crawlers discover your XML sitemap independently of any submission. The checker reports every Sitemap line it finds so you can confirm it is declared.

Is robots.txt a security or access-control tool?

No. Robots.txt is a public, voluntary instruction for cooperating crawlers, not access control. Anyone can read it, and it cannot protect private content. Use authentication or server rules to actually restrict access.

Does this checker simulate exactly how Googlebot reads my file?

It parses and reports your directives as written, but it does not replicate any single crawler's full internal matching policy. Treat the grouped rules and counts as an accurate readout, not a crawler-specific simulation.

Nhiều kiểm tra miễn phí hơn

Khám phá tất cả Công cụ Miễn phí của Novaverb

Trình kiểm tra SEO trang webPhạm vi thu thập thông tin, các trang có thể lập chỉ mục và liên kết nội bộ
Keyword Research ToolFree AI keyword research tool for search volume, keyword difficulty, …
SERP CheckerCheck live organic search results and AI Overview rankings for …
Website Security CheckerAudit website security posture, TLS/SSL certificates, HTTP security headers, and …
Server Response Time CheckerMeasure server Time to First Byte (TTFB), DNS lookup, TCP …
Backlink CheckerExplore backlinks, referring domains, dofollow links, and domain authority for …
Sitemap CheckerValidate XML sitemap structure, URL counts, reachability, and index type. …
Meta Tag CheckerCheck page title length, meta description, H1 heading structure, Open …
HTTP Status & Redirect CheckerTrace HTTP status codes (200, 301, 302, 404, 500) and …
Website MonitorMonitor website availability, HTTP status code, and server response time …
HTTP/2 TestCheck whether your web server supports HTTP/2 via TLS ALPN …
HTTP/3 TestTest whether your web server supports HTTP/3 over QUIC with …
Website Performance TestTest global website loading speed and TTFB waterfall timing across …
GEO CheckerAudit whether ChatGPT, Perplexity, Gemini, and Google AI Overviews can …
Core Web Vitals CheckerCheck real-user Core Web Vitals (LCP, INP, CLS) from Chrome …
PageSpeed CheckerRun a live Lighthouse performance audit to test PageSpeed, Core …
Website Safety CheckerCheck whether a domain or URL is flagged for malware, …
Knowledge Graph CheckerCheck whether a brand, person, product or organization is recognized …
Keyword Gap CheckerCompare your site against competitors to find missing high-traffic keywords, …
Competitor Top PagesFind top organic traffic-driving pages for any competitor domain. See …
Duyệt trung tâm công cụ miễn phí đầy đủ
Kiểm tra → hiểu → sửa

Biến kiểm tra này thành một sửa chữa đã được xác minh

Mỗi công cụ miễn phí của Novaverb là một kênh: thực hiện kiểm tra, hiểu bằng chứng, sau đó sửa chữa và chứng minh rằng nó đã được giải quyết với một lần kiểm tra lại mới - không có trạng thái vượt qua nào được tạo ra.