Noindex a PDF with X-Robots-Tag: Check the File Response
Keep a public PDF downloadable while excluding it from search using response headers, with a verification checklist and examples.
A robots meta tag on the page linking to a PDF does not apply to the PDF itself. For a non-HTML resource, inspect the response from the file URL and configure the appropriate HTTP header there.
Google documents X-Robots-Tag for directives on non-HTML resources. A public PDF with noindex remains publicly accessible; use authentication or remove the file if access itself must be restricted.
Identify the resource you mean
In this fictional example, three URLs serve different jobs:
| URL | Role | What to inspect |
|---|---|---|
/reports/annual-report | HTML landing page | Its own meta directives and canonical |
/downloads/report-2025.pdf | The PDF download | Its HTTP response headers |
https://files.example.com/report-2025.pdf | A redirected file host | Final response and host configuration |
Noindexing the HTML page does not remove the PDF. A rule on your main domain may not affect a file served from another hostname.
The intended response
This is an illustrative response, not an instruction for a particular hosting provider:
HTTP/2 200
Content-Type: application/pdf
X-Robots-Tag: noindex
Configure the header through your server, hosting platform or CDN for the intended file/path. Start with one non-sensitive test document. Do not apply a sitewide noindex header merely to exclude a single file.
Check with a GET request
A HEAD request is useful, but some servers handle HEAD and GET differently. Save the headers from an actual GET when confirming the download:
curl -sS -L -D pdf-headers.txt -o downloaded-report.pdf 'https://example.com/downloads/report-2025.pdf'
Read the last response after any redirects. Confirm the final URL, status, content type and X-Robots-Tag. The output file should be the expected PDF, not an HTML access/error page. This checks the response to your request; it does not establish what Google received earlier.
Why the PDF can remain visible after a change
| Observation | Investigation |
|---|---|
| Only the attachment/landing page has noindex | Apply and inspect the directive on the PDF response |
| Initial redirect has a header; final PDF does not | Check the final host and response configuration |
| Cached response lacks the header | Check the relevant CDN/cache rule and retest |
| robots.txt blocks the PDF | Google may be unable to read the noindex header; review the crawl policy |
| Current file is correct; indexed evidence is older | Record the change and allow recrawl/processing |
| Another copy appears in results | Inspect that exact URL separately |
Do not remove a crawl block blindly from a sensitive file: make the access decision first. For public resources intended only to leave search, Google must be able to fetch the directive. See noindex versus disallow.
A copyable verification record
Requested file URL:
Final file URL:
Checked at (UTC):
GET status and content type:
X-Robots-Tag value:
Robots access for the file:
Where the header is configured:
Search Console last crawl/date evidence:
Other copies checked:
Use the directive tool to interpret a copied value; it does not fetch or audit PDFs. For temporary search suppression while arranging a permanent fix, consult Google's removals guidance. A removal request and a permanent indexing policy are different actions.
If the file is meant to consolidate with an equivalent HTML version instead, investigate a canonical HTTP header rather than assuming noindex is the right goal. Preserve a dated response sample and recheck the exact resource after changes.