Crawl Explorer: the SEO crawler that keeps every captured URL and its evidence
Kiểm tra từng URL đã thu thập trong lượt thu thập dữ liệu của dự án, và biến mỗi phát hiện kỹ thuật có bằng chứng thành một hành động.
What is an SEO crawler, and what is Crawl Explorer?
An SEO crawler is software that fetches a website's pages the way a search engine's bot does and records what each URL returned: status code, canonical, indexability signals, titles, headings, depth, internal links and rendering. Crawl Explorer is Novaverb's SEO crawler, and its defining habit is that it keeps the complete captured set as a timestamped record, so a finding is always a list of pages you can open and compare.
Crawl Explorer trình bày mọi URL được thu thập trong giới hạn phát hiện, quyền và kế hoạch của dự án, bao gồm trạng thái, độ sâu, siêu dữ liệu và tín hiệu sức khỏe. Tập hợp kết quả đã lưu không được lấy mẫu, và cùng một lần thu thập được sử dụng cho phần còn lại của Novaverb.
What does the SEO crawler capture for each URL?
The SEO crawler captures, for every URL within the project crawl's discovery, permission and plan limits, the fields a technical review needs: HTTP status and redirect chain, declared canonical and the status of what it points at, indexability directives, title, description and H1 state, crawl depth, internal links in and out, and render evidence where JavaScript rendering ran. Nothing is sampled.
- Every stored crawl URL with status, canonical, indexability, depth and internal links
- Duplicate clusters for titles, descriptions and near-identical bodies when the fields are captured
- Render evidence: what the page looked like after JavaScript, where rendering ran
- Data-quality states that say which fields were captured and which were not
- Filters and exports on supported fields, so a finding becomes a list you can work
- A timestamped record of each crawl you can compare with the next one
Who is the SEO crawler for?
The SEO crawler is for people who are asked to prove a technical problem, not just to suspect one: a technical SEO working a large site with more URLs than a spreadsheet holds, an agency that must show a client which pages were affected, a developer who needs the exact URLs behind a ticket, and an in-house lead comparing this month's crawl with last month's.
Technical SEOs
Interrogate a crawl one question at a time: status class, redirects, canonicals, duplicate titles, H1 state, render risk.
Đại lý
Show a client the list of affected URLs behind every finding, with the crawl date it came from.
Các nhà phát triển
Take the exact URLs and captured fields into a ticket, then re-crawl to confirm the change landed.
In-house leads
Keep every crawl as a record and compare scans over time instead of trusting a stale export.
Why use an SEO crawler that keeps its evidence?
Use an SEO crawler that keeps its evidence because a technical finding without the pages behind it is an opinion. Crawl Explorer stores every URL it fetched with the signals it captured, so a duplicate-title finding is the list of duplicated pages, a canonical problem shows what the canonical points at, and a page that was never fetched is never presented as a page that passed.
Tập hợp đã ghi lại hoàn chỉnh
Kiểm tra từng URL đã lưu từ lần thu thập dữ liệu cùng trạng thái HTTP đã thu thập, canonical đã khai báo, tín hiệu khả năng lập chỉ mục, độ sâu và liên kết nội bộ, sau đó lọc hoặc xuất các trường được hỗ trợ.
Các phát hiện có bằng chứng
Các cụm trùng lặp, bằng chứng render và trạng thái chất lượng dữ liệu được hiển thị khi các trường thu thập dữ liệu cần thiết có sẵn.
Bằng chứng bạn sở hữu
Mỗi lần thu thập dữ liệu là một bản ghi có dấu thời gian mà bạn giữ lại; bạn có thể so sánh các lần quét theo thời gian và dùng một lần chạy detector sau đó để kiểm tra xem một vấn đề được hỗ trợ có còn hay không.
How does the SEO crawler work, from crawl to fix?
The SEO crawler works in four steps inside your project. You start a crawl within the project's discovery, permission and plan limits; Crawl Explorer stores every URL it retrieved with the fields it captured; you open the cockpit that matches your question and filter down to the affected URLs; then you fix or export them and re-crawl to compare, because the earlier list is not assumed to still be true.
Crawl the project
Start a crawl of the site within the project's discovery, permission and plan limits. Robots directives and your settings decide what is fetched; nothing outside them is guessed.
Read the captured set
Every retrieved URL is stored with status, canonical, indexability, titles, headings, depth, links and render evidence, plus a data-quality state for fields that were not captured.
Open the cockpit for your question
Eighteen focused views each ask one thing of the crawl, from redirect health to duplicate clusters, and filter to the URLs affected.
Fix, export, re-crawl
Work the list, export the supported fields to a ticket or a sheet, then run the next crawl and compare, so a fix is closed by new evidence.
The SEO crawler stores the complete captured set, never a sample
Every URL retrieved within the crawl's limits is kept with the fields captured for it: HTTP status and redirect chain, declared canonical and its target's status, indexability directives, title, description, H1, crawl depth, internal links in and out. A finding filters this set rather than estimating from part of it, so the count on a finding is the count of pages you can open.
- Status and redirect chain per URL
- Canonical and its target's status
- Depth and internal links in and out
The SEO crawler shows render evidence and data-quality states
Where JavaScript rendering ran, the crawler keeps what the page looked like after rendering next to the raw response, so a title or a link that exists only after scripts run is visible as such. Where a field could not be captured, the data-quality state says so; a missing signal is reported as missing, not as a pass or a zero.
- Render evidence where rendering ran
- Duplicate clusters from captured fields
- Unavailable fields stay unavailable
The SEO crawler keeps every crawl as a record you can compare
Each crawl is a timestamped record you own. Compare scans over time, use a later detector run to check whether a supported issue remains, and let the same captured crawl feed the rest of Novaverb, from Site Health Audit's grading to the internal link graph in Backlinks Explorer, so every pillar reads one crawl, not its own copy.
- Timestamped crawl history
- Recheck against a fresh crawl
- One crawl feeds every pillar
An evidence-keeping SEO crawler vs. a sampled site scan
A sampled site scan fetches a slice of the site, extrapolates a score and asks you to trust the extrapolation; when you ask which pages, it cannot say. An evidence-keeping SEO crawler stores everything it fetched, filters that set for each finding and keeps the crawl the finding came from. One gives you a number about the site; the other gives you the pages.
- Fetches a slice and extrapolates a site-wide score
- Cannot list the pages behind a finding
- Treats an unfetched page as a pass or a zero
- Overwrites the last scan, so change is assumed
- Stores the complete captured set within the crawl's limits
- Every finding is a filtered list of URLs you can open
- Unavailable fields are reported as unavailable
- Every crawl is a timestamped record you compare
What data does the SEO crawler produce and where does it come from?
The SEO crawler reads one source: your own crawl of your own site. Every title, canonical, directive, status code and rendering signal it shows was retrieved from your pages, not modeled or estimated from a third party. The crawl runs within the project's discovery, permission and plan limits, respects robots directives, and records for each field whether it was captured.
The project crawl
Fetched pages with their HTTP response, headers, HTML, canonical, directives, titles, headings and links.
Render evidence
Where rendering ran, the page after JavaScript, kept beside the raw response for comparison.
Crawl limits and directives
Robots rules, discovery scope and plan limits decide what is fetched; the crawl records what they excluded.
Earlier crawls
Each crawl is stored, so a later run is compared with the record instead of replacing it.
What can you produce with the SEO crawler?
What you can produce with the SEO crawler is the material a technical fix needs: a filtered list of affected URLs for each cockpit, exports of supported fields for tickets and sheets, duplicate clusters, canonical maps, redirect chains, render comparisons, and a before-and-after view when the next crawl lands. Each output names the crawl it came from.
How is Novaverb's SEO crawler different?
Novaverb's SEO crawler is different in what it refuses to invent. It does not sample and extrapolate, it does not turn an unfetched page into a pass, it does not overwrite the last crawl, and it does not keep a private copy of the data for each tool: Site Health Audit, Backlinks Explorer and the content studio all read the same captured crawl.
How does the SEO crawler help technical SEO?
For technical SEO, the SEO crawler is where the evidence starts. It shows which URLs return errors or sit in redirect chains, which canonicals point at pages that are not indexable, which titles and descriptions are duplicated, which pages have no H1 or are buried too deep, and which links point at pages that no longer answer. Each of those is a list of URLs, with the crawl date attached, that you can work, export and re-check.
How does the SEO crawler help AI answer visibility?
For AI answer visibility, the SEO crawler supplies the page structure the readiness checks read: headings, question coverage, structured data and rendered content. GEO & AEO Readiness grades that captured structure, and the content studio writes from it, so a page can only be prepared for an answer engine to the extent the crawl saw it. Readiness measured here is on-page readiness; whether an engine cites the page is observed separately.
Which plans include the SEO crawler?
The SEO crawler is included from the Free plan, so a first crawl needs no card. What changes between tiers is scale: how many URLs a crawl may capture, how many audits and data credits a month, how many seats and workspaces, and how much crawl history stays open for comparison. The crawler itself behaves the same on every plan.
Hoạt động với
How much does the SEO crawler cost?
The SEO crawler costs what your plan costs; it is not licensed separately, and the Free plan includes a first crawl. Plans are priced on scale, URLs per crawl, audits and data credits a month, seats, workspaces and history depth, and the current amounts for each tier live on the pricing page, read from one source so this page never quotes a stale number.
So sánh các gói →Best practices for running an SEO crawler
The best practices for running an SEO crawler are about scope and rhythm. Set the discovery scope so the crawl covers the site you are responsible for, let robots directives stand, crawl before and after every structural change, read the data-quality state before trusting an absence, and keep the crawl history rather than exporting once and forgetting.
- Scope the crawl to the site and sections you are accountable for
- Crawl before and after a migration, a template change or a redirect rollout
- Read the data-quality state before treating a missing field as a real absence
- Open the cockpit that matches the question instead of scrolling the whole set
- Compare scans in the crawler rather than in a spreadsheet you exported last month
How does the SEO crawler feed the rest of Novaverb?
The SEO crawler feeds the rest of Novaverb because every pillar reads the same captured crawl. Site Health Audit grades it against NovaBrain criteria, Backlinks Explorer draws the internal link graph from it, GEO & AEO Readiness checks its structure, the content studio writes from its pages, and Reports carry its findings, dated, to the people who asked.
Frequently asked questions about the SEO crawler
How many URLs can the SEO crawler capture?
The crawler captures every URL it can reach within the project's discovery scope, its robots and permission rules and the URL limit of your plan. The limit is a plan number, not a sampling rate: up to that limit the set is complete, and the crawl records what the limit or the directives excluded.
Does the SEO crawler render JavaScript?
Where rendering runs for a project, Crawl Explorer keeps the page after JavaScript beside the raw response and reports render evidence, so content or links that appear only after scripts execute are visible as such. Where rendering did not run, the render fields are marked unavailable rather than assumed.
Can I export what the SEO crawler found?
Yes. Every cockpit filters the stored set to the affected URLs and lets you export the supported fields, so a finding travels into a ticket or a spreadsheet as the list of pages and values behind it. The crawl date goes with it, so the export says which scan it describes.
How does the SEO crawler tell a fixed issue from an unfixed one?
By re-crawling. A later crawl is a new timestamped record, and a later detector run checks whether a supported issue is still present on the fresh evidence. A fix is closed by new evidence, not by someone marking the finding done, and the earlier crawl stays available for comparison.
Does the SEO crawler respect robots.txt?
Yes. The crawl runs within the site's robots directives, the project's discovery scope and your own settings. Pages those rules exclude are not fetched and are not presented as if they had been; the crawl records the exclusion so an absence in the set has an explanation.
Is the crawl data shared with Novaverb's other tools?
The same captured crawl feeds every pillar. Site Health Audit grades it, Backlinks Explorer draws the internal link graph from it, GEO & AEO Readiness reads its page structure and the content studio writes from its pages, so findings in one tool agree with the pages another tool shows.
Run your first crawl with the SEO crawler
Run your first crawl with the SEO crawler by creating a free workspace, adding your site as a project and starting a crawl; the first captured set is the evidence every other pillar builds on. Compare the plans if you already know the URL volume and crawl history you need.