Theos Group

How findable is your business for Google and for AI? Free, instant.

Prefer the full report right away? Comprehensive scan — € 49 →

Knowledge base

Can Google actually see your website?

A small technical file, robots.txt, determines whether Google may view and render your site, and an error in it can halt your entire visibility without you noticing anything.

What is that file anyway

Every website has, or should have, a file named robots.txt. You find it by typing /robots.txt after your domain name, for example yourbusiness.nl/robots.txt. This file is the first thing Google reads before it views your site. It is in plain text and tells search engines which parts of the site they may and may not visit.

You can compare it to a sign at the entrance of a building. Some doors are open, others are closed with a simple note. The problem is that businesses sometimes hang the wrong note by accident. Instead of closing off one back door, they close the entire building to visitors, including the most important rooms.

Beyond the paths that may or may not be visited, there is also something more subtle: the files that build the page. Think of CSS, which handles formatting, and JavaScript, which handles movement and interaction on the page. If robots.txt blocks those files, Google can visit the page, but cannot see it properly. It then sees a bare, broken version of your site.

This is closely linked to another checkpoint: whether your core content is accessible without JavaScript, as we explain at core content is accessible without JavaScript. Both points address the same question: can Google see your page as a visitor sees it, or does Google get an incomplete picture that scores poorly.

Why this is more than a technical detail

Google uses a process called rendering: the search engine essentially rebuilds your page, just as a browser does, to see what a visitor actually gets to see. This is only possible if Google is allowed access to all components needed for that. If CSS or JavaScript is blocked, Google sees a bare page without formatting, without menu, sometimes without text that is loaded by scripts.

The result is that Google misjudges what your page is about. A page that looks fine to visitors can seem incomprehensible or empty to Google. This is not a matter of scoring slightly worse. It is a matter of not being included, even though you did nothing wrong with the content itself.

There is a difference between a page that does not exist and a page that exists but has been made inaccessible. This also touches on other points in our checklist, such as whether a removed page properly returns a 404 status, as we discuss at the 404 page itself is helpful. In both cases, it is about clarity: show Google what is and is not meant to be found.

How to check if everything is in order for you

Type your domain name followed by /robots.txt in your browser's address bar. A simple list of rules appears, usually with words like 'Disallow' followed by a path. Disallow means: this path may not be visited. Check if there is anything there you do not expect, such as a slash alone, which means the entire site is blocked.

Pay special attention to rules that refer to folders with names like /wp-content/, /assets/, /css/ or /js/. These are often precisely the folders where your site's formatting and interaction are stored. If those are blocked, there is a good chance Google cannot properly render your pages, even if the rest of the file seems harmless.

Google Search Console, a free service from Google for website owners, has a test with which you can check a specific page for rendering problems. You see there an example of how Google perceives your page. If that looks bare or broken while the page looks fine in your browser, that is a clear signal that something is going wrong with the loading of CSS, JavaScript or images.

If you are unsure whether a particular path has been deliberately excluded or accidentally blocked, consult with whoever built your website. Some blocks are intentional and sensible, such as excluding an internal search result—a topic we address separately under internal search result pages on noindex. Other blocks are simply errors from a default setting that was never reviewed.

What it costs if this is not in order

If robots.txt blocks important paths, you inadvertently exclude parts of your site from Google. These could be product pages, blog articles, or even your entire site after migrating to a new system. This happens more often than people think, especially after a website redesign, when a setting intended for the test environment accidentally remains on the live site.

The insidious part is that you notice nothing in your day-to-day use of the site. The pages load normally, and visitors who receive the link directly can still access them. Only traffic from search engines slowly dries up, without any obvious reason. Many entrepreneurs then look for the cause in content or competition, when the problem actually lies in a technical file they have never opened.

When only CSS and JavaScript are blocked, the effect is more subtle but no less damaging. Your pages remain discoverable, but Google assesses them based on an incomplete picture. Text loaded by a script may not count. A page that looks perfectly fine to you may be rated by Google as thin on content, with all the consequences for your position in search results.

Recovery is usually not a matter of months of work. Often it comes down to adjusting a few lines in a file. The point is that you must first recognize the problem. Without inspection, a block sometimes remains unnoticed for years, simply because nobody looked at it after the site was last modified.

What needs to happen to put this in order

The starting point is simple: robots.txt should only block paths you deliberately want to exclude, such as an admin login page or a shopping cart that should never appear in search results. Everything that contributes to how the page looks or functions must remain accessible. That is the entire principle behind this checkpoint: do not lock things down as a precaution, but block deliberately where necessary.

Have this file checked as part of a broader review of what should and should not be indexed. That review—an indexation matrix per page type—maps out which pages should be visible to Google and which deliberately should not, as we explain under indexation matrix per page type. Robots.txt is then the tool to execute those decisions, not the place where those decisions are made.

Check this file not only when building the site, but also after every major change: a migration to another system, a new website builder, or a restructuring of the site. These are precisely the moments when settings are accidentally carried over from a test environment. A brief check afterwards prevents a temporary setting from becoming permanent.

Also keep in mind that robots.txt does not stand alone. It is closely linked to other signals you send to Google, such as the page-level instruction via a noindex tag, which we cover more fully under noindex via meta robots or the X-Robots-Tag. Both must not contradict each other. A page you want to show must not be opened via one channel and still blocked via another.

Frequently asked questions

Can I modify robots.txt myself without a website builder

Technically, it's a simple text file that you can open and modify yourself. In practice, caution is advisable: a single incorrect line can shut down your entire site to Google. Have any change reviewed by whoever manages your site, or test the modification first with the tool in Google Search Console before you implement it permanently.

Does a robots.txt block mean the page immediately disappears from Google

Not immediately, and not always completely. The effect often builds gradually as Google encounters the block more frequently. Sometimes a page remains visible without a description, because Google knows the title but is not allowed to view the content. In all cases, the result is a weaker position than if the page were fully accessible.

I don't have a robots.txt file, is that a problem

Without this file, Google is in principle allowed to go anywhere, so in that sense there is no direct block. Still, it is advisable to have one, if only to deliberately exclude a few paths, such as an internal search system. This also aligns with checkpoint 1.2, where we go into the precise configuration of this, as can be read at checkpoint 1.2.

Further reading

How is your own website doing?

Two figures, within a minute, free of charge — and you don't need to leave anything behind.

How findable is your business for Google and for AI? Free, instant.

Prefer the full report right away? Comprehensive scan — € 49 →

Want to know what's causing it?

This scan gives you the status. The report gives you the causes and the order in which you address them.

  • For each finding, what's wrong and what needs to happenThe free scan gives you the status. The report gives you the list — in plain language, ranked by how much it matters, without you needing to understand anything technical.
  • All four AI assistants instead of oneThe free result is a sample from one assistant. The report asks eight questions of all four, so you know whether the issue is with one assistant or all of them.
  • Which companies do get mentionedWho gets the customer you're losing? Those names are in the report, with how often they appear where you're missing.
  • A deeper measurement of your websiteThe free scan looks at ten pages. The report goes through up to fifty, including the pages where your customers ultimately end up.