How do search engines work? Indexing, crawling, and ranking Post-24

How do search engines work? Indexing, crawling, and ranking Post-24

How do search engines work? Indexing, crawling, and ranking Post-24

How do search engines work? Indexing, crawling, and ranking


It goes without saying that search engines have made our lives easier with various features. From checking the weather to setting mobile alarms, search engines are used everywhere. People all over the world perform an average of 3.8 million searches per minute on Google alone. In a day, it goes up to 5.6 billion. The amount of traffic and challenges a search engine has to face every day can be estimated from the statistics of Search Tribune. But have we ever thought about how this search engine works?

 

In today’s post, we will try to find out, “How search engines work”? But before that, for the convenience of readers, we will discuss some important terms that they often do not understand. Our writing topics are beautifully arranged. If you already have knowledge about a topic, then you can easily skip it. So I hope there will be no reason to get bored. We write content considering each of our audiences. Everyone considers Google to be the universal search engine. So Google will dominate our discussion as well.

How do search engines work?

How do search engines work?

Search engines have three primary functions. They are- crawl, index, and rank. First, the crawling system uses to find information, it goes to places that are already indexed. Then Googlebot determines who should be shown first and who should be shown later. It is good to clear up another thing, for the convenience of the reader. Each search engine has its own crawling robots. They rank the content by considering the previously specified ranking factors. One such bot is Google Bot. For example, Moz.com has Rogerbot, Bing has Bingbot.

 

We can imagine that some terms are confusing to the reader. Therefore, we will return to a detailed discussion of all of them.

 

What is search engine crawling?

What is search engine crawling?

The direct Bengali meaning of crawl is – to crawl. In the case of search engines, the word means much the same, that is, when you search for a keyword, Google first searches for the results that are already indexed for you. This search process is called crawling. And the one who crawls is called a ‘crawler’, in many cases it is called a ‘spider’. The one who crawls on behalf of Google is called a ‘Google bot’. Crawling can be anything. It depends on your search.

 

Google bot first targets some web pages by fetching (a method of quick search). Especially those whose URLs are completely new. They are first identified and stored in Caffeine. Caffeine is a huge URL repository where new URLs are stored. Google Caffeine was created in 2010. Let’s leave that story for another day. I don’t want to bore the reader by saying it all together. By saving their URLs in advance in Caffeine, it does not take much time to show their search results. For this, they are crawled in advance and saved in Caffeine.

 

What is search engine indexing?

Search engine indexing is the process of search engines storing their search results in advance for the convenience of showing them. It is a bit like arranging books in a library. Just as it is convenient to know where a book is if you arrange it in advance, it is very convenient for search engines to show results later if they index it in advance. For this reason, every search engine indexes it in advance. I hope the reader understands what search engine indexing is.

 

What is search engine ranking?

When a user searches for a keyword, the search engine refines and enhances it according to the keyword and provides the most relevant data. Google and other search engines have several rules and regulations that they follow to show your content first. And these rules and regulations are called ranking factors. Usually Google has about 217 ranking factors. However, it is very variable. Sometimes less and sometimes more. However, in my case, from the most advantageous position, it can be said that there are more than 200 ranking factors. And after going through this method, the whole process of who will show first and who will show gradually at the end is controlled by this is called ranking factors. Another thing to say in this case is that-

 

The earlier the search engines show the results, the more relevant Google thinks the result is for you.

 

In one of our next posts, we will discuss some of the top ranking factors. For those who are more interested in SEO, something good is waiting for them.

 

How do search engines organize information?

 

Search engines index and prepare information for you when you search for something. For the convenience of all of us, we have arranged today’s post by following Google. That is, we have shared information accordingly by taking Google as the search engine. Although I have said it before, I am saying it again. Many people may have come here directly without reading from the beginning. I have informed them again for their benefit. According to Google’s information, they crawl more than 100 billion web pages and organize the search results for you in advance.

 

Basics of Search

The crawling process starts from the previously listed web address and through the sitemap provided by the website owners. For those who are new, I think it is necessary to add another piece of information. Sitemap is like a file manager. The file manager on your mobile is set up – where are the video files, where are the audio files. Similarly, the name of the website file is sitemap. Through it, search engines get a fair idea of ​​where and what data is on your website, what posts are there. Which is later shown to the search engines by matching them according to keywords.

 

If the reader pays a little attention, one of the ‘means’ for search engines to know about your website or web pages is – sitemap. Through this sitemap, search engines are able to have an idea about your website. The main purpose of talking about all this is – this sitemap is the only way to know if any information on your website has been updated.

 

After that, the computer program determines what information the search engines will find. Accordingly, they take the information of your website. Now if you say that my website has some very important information. Which I do not want the search engine to read. In that case, what should I do? If this kind of question has come to your mind, then thank you sincerely. Now let’s come to your answer. Yes, the search engines have definitely set some rules for you, the reader. They have named it – ‘robots.txt’. Through this ‘robots.txt’, you will be in complete control of your website. That is, you can tell the search engines in advance which web pages can be searched and which web pages cannot be searched.

 

Since we have kept Google as the search engine for today’s discussion – it is good to mention another thing. That is, the search console, through which you will get some more important control of your website.

 

For example, suppose you have made some new changes to your website. Or you have moved some web pages to another place. In this case, the problem is that since Google has already indexed a structure of your website, now when it needs to show this same result, Google will not find the previous result. There are two reasons for this. First, we have already discussed how Google shows results from the index made earlier or through the sitemap. But when you have changed the structure of your sitemap, Google no longer knows about it. Second, you have not informed Google that you have changed your website. That is, Google does not know anything about the new sitemap. You have not asked Google to re-crawl.

 

The solution to all these problems is to inform Google about any changes to your website through sitemap updates. Or if there is no fundamental change in the structure of the site and everything is tidy, then it does not cause any problems. And in order to control this entire process, you can resort to ‘Google Search Console’. Another thing is good to know – Google does not charge you any extra money for re-crawling the sitemap repeatedly.

 

Finding information through crawling

The number of web pages in the world is constantly increasing. According to the information of Siteify, an average of 54,7200 web pages are being created every day. You can understand, readers, how much the rush to create web pages is increasing. It seems like – the number of books in libraries is increasing in large numbers. Think for a moment, readers – if suddenly a large number of books are sent to a library, then the librarian does not supervise them or does not register them according to the ISBN number, how much chaos will be created? It would be wise not to think about that in vain. But Google is not like the librarians of the government libraries of our country. They index the right information at the right time unless there is a block from the site. They have named the software they have created for this – “Web Crawlers” whose job is to index the constantly created publicly accessible web pages. You will understand better why I said publicly accessible when I discuss ‘robots.txt’.

 

I cannot help but add something more in this regard. When Google scrolls a specific web page of your website, it does not just scroll that specific web page and comes back. Rather, it scrolls all the links, backlinks, and inbound links added to that web page. Yes, but of course those that are approved by ‘robots.txt’. In this way, they fill their database by scrolling new web pages with their “web crawlers”.

 

Organizing and indexing information

Let’s take a look at what Google has to say about this-

 

When crawlers find a webpage, our systems render the content of the page, just as a browser does. We take note of key signals — from keywords to website freshness — and we keep track of it all in the Search index.

 

Here’s an important piece of information to note: ‘We take note of key signals’ means that they are collecting some comments about your webpage. And they also say that in those notes, they are taking into account everything from keywords to the freshness of the website.

 

[/copy]