What a search engine reads on a page
Every page carries a second set of facts, written for machines: title, description, canonical, robots, headings, links. What each one does, in Google’s words.
A page is read twice. A person reads what is on the screen. A search engine reads the document underneath: the same words, and a few lines that were written for it and that a visitor never sees.
Below is each of those lines, as it stands on our own home page, with what Google’s documentation says it is for. We quote Google instead of summarising it. Most of what is repeated about these lines is stronger than what Google actually says.
The title
<title>TULUS — Web development and SEO studio in Bandung</title>
This is the line you click in a search result. It is a suggestion. Google builds that line from
“a number of different sources”, and the <title> element is the first on
its list. It may use another:
If we’ve detected an issue on the page, we may try to generate an improved title link from anchors, on-page text, or other sources.
So we write one title for each page, and make it say what that page is.
The description
<meta name="description" content="TULUS is a web development and SEO studio in Bandung, Indonesia. Websites people understand and search engines can read, with the structure shown underneath.">
The few lines under the title of a search result. This one is an offer too. In Google’s words:
Google sometimes uses the meta description HTML element if it might give users a more accurate description of the page than content taken directly from the page.
And the text is not fixed: “Google Search might show different snippets for different searches.” What you can do is write a plain summary of the page and leave it there.
The canonical address
<link rel="canonical" href="https://tulus.studio/">
One page can usually be reached at several addresses: with www and without, with a slash at the
end, with a tracking code after a question mark. This line names the one you mean. Google calls it
a “strong signal”, and also says that
“indicating a canonical preference is a hint, not a rule.”
That is why the line is not enough by itself. On this site the other spellings of an address do not show the page at all. They redirect to the one address.
Robots
<meta name="robots" content="index, follow, max-image-preview:large">
This says whether the page may appear in search results. The word that matters is the opposite one,
noindex, and it has a trap. A page must be fetched before the line can be read.
Google’s documentation:
If the page is blocked by a robots.txt file or the crawler can’t access the page, the crawler will never see the
noindexrule, and the page can still appear in search results, for example if other pages link to it.
So a robots.txt file is the wrong tool for keeping a page out of Google. Google says of robots.txt that “it is not a mechanism for keeping a web page out of Google.”
The language
<html lang="en">
Here the two readers part. Google does not use this line. It says so:
Google uses the visible content of your page to determine its language. We don’t use any code-level language information such as
langattributes, or the URL.
A screen reader does use it. The W3C gives the reason in one sentence: “Screen readers can load the correct pronunciation rules.” The line stays on every page, for the person listening.
Headings
The same again. From Google’s starter guide, in a list of things it believes you should not focus on:
Having your headings in semantic order is fantastic for screen readers, but from Google Search perspective, it doesn’t matter if you’re using them out of order.
That sentence is about order and number. It does not say headings are ignored: headings are on the
list of sources for the title above. We keep one h1 and an unbroken order on every page anyway. A
screen reader lets a person move through a page heading by heading, and the order is their map.
Links
<a href="/work">Work</a>
A link is how a search engine gets from one page to the next, and Google is exact about what counts as one:
Generally, Google can only crawl your link if it’s an
<a>HTML element (also known as anchor element) with anhrefattribute.
A button that changes the address with a script looks like a link and is not one. The words inside the link count as well: “Good anchor text is descriptive, reasonably concise, and relevant to the page that it’s on and to the page it links to.” The words “read more” are none of the three.
Pictures
<img src="…" alt="Home page of laras.tech: the headline “Make your business intelligent.” beside a dark 3D scene of a desk with lamps, a parcel, a phone and a screen.">
The alt text is what is read aloud in place of a picture, and what is shown when the picture does
not arrive. Google reads it too:
“Google uses alt text along with computer vision algorithms and the contents of the page to
understand the subject matter of the image.”
Structured data
<script type="application/ld+json">{ "@type": "Organization", "name": "TULUS", … }</script>
A block of facts in a form made for machines. On our home page it says that TULUS is an organisation, that this is its website, and that this is a page of it. It makes a page eligible for the richer kinds of search result, and no more. Google:
Using structured data enables a feature to be present, it does not guarantee that it will be present.
We state what is true and stop. There is no rating in ours, because nobody has rated us.
The words themselves
All of this assumes the text is in the document. When a page is put together in the browser by a script, a search engine has to run the script before there is anything to read. Google does run it, after a wait: “The page may stay on this queue for a few seconds, but it can take longer than that.” The same page goes on:
Keep in mind that server-side or pre-rendering is still a great idea because it makes your website faster for users and crawlers, and not all bots can run JavaScript.
Every page of this site is complete before any script runs.
Read your own
Open your home page, right-click, and choose “View page source”. Search for <title, then
description, canonical, robots. It takes five minutes, and it is what the search engine was
given.
Or press X-ray this website on our home page. It reads these lines from the page you are on and sets them out beside it.
What this list is not
It is not a way to rank. Nothing above makes a page worth finding. These lines make sure that a page which is worth finding can be read, and is described the way you meant.