-
All about Google-Must Read
How Do I Get My Site Listed on Google?
1. The basics
Google is a fully automated search engine that employs robots known as "spiders" to crawl the web and find sites for inclusion in the Google index. Since this process doesn't involve human editors, it's NOT necessary to submit your site to Google in order to be included in our index. In fact, the vast majority of sites listed aren't manually submitted for inclusion.
Google doesn't accept payment for inclusion (known as "paid inclusion") of sites in our index, nor for improving the rank of sites in our results. We do offer advertising opportunities adjacent to our results, which are always clearly labeled "Sponsored Links." The method by which we find pages and rank them as search results is determined by many different factors, including the PageRank technology developed by our founders, Larry Page and Sergey Brin.
2. Submitting a site
We add thousands of new sites to our index each time we crawl the web, but you may submit your URL as well. Submission isn't necessary and does not guarantee inclusion in our index. Given the large number of sites submitting URLs, it's likely that your pages will be found in an automatic crawl before they make it into our index through the URL submission form. We DO NOT add all submitted URLs to our index and cannot predict when or if they will appear.
Please visit our Add URL page to input your URLs. You can submit your site as often as you like, but multiple submissions won't improve the likelihood of your site being added or accelerate the process. We don't penalize sites for "over-submitting." If you choose to submit your site, only the top-level domain is necessary, as the spiders can follow your internal links to the rest of the pages.
You may also use the Google Sitemaps (Beta) program to create and submit a detailed sitemap of your pages. We're testing this as a complement to our current crawl and encourage webmasters to participate. Google Sitemaps makes it easier for webmasters to provide information about their sites, and to update us when pages are added or changed.
The best way to ensure that Google finds your site is to have pages on other relevant sites to link to yours. Google's robots jump from page to page on the web via hyperlinks, so the more sites that link to your pages, the more likely it is that we'll find them quickly.
-
2
My webpages have never been included in the Google index.
1. My site's new to the web, and I recently submitted it.
Google finds sites through a process known as "crawling" the web. This involves robot software that follows hyperlinks from site to site. Google currently looks at billions of URLs during our crawls.
When a URL is submitted to Google, we're able to look for it in our next crawl. If you've already submitted your URL, your site could easily appear in our search results after our next crawl. However, if no other sites links to yours, it may be difficult for our crawler to find you. Conversely, if many sites link to your page, there's a good chance we'll find you even without the submission of your URL.
2. My site's been live for a few months.
If we haven't picked up your site after several months, it's possible that our spiders aren't able to find your pages. If you increase the links pointing to these pages, it'll improve the chance that we'll find your site.
It's also possible that we're not able to crawl your site due to technical reasons. A few of the most common ones are listed below:
Your pages were unavailable when we tried to crawl them. If a page is down due to network or hosting problems, we try to visit it multiple times. If our crawlers can't reach it, it won't be listed in our index. If such unavailability was a transient problem, your site will likely be added to our index again soon.
Your pages are dynamically generated. We're able to index dynamically generated pages. However, because our web crawler could overwhelm and crash sites that serve dynamic content, we limit the number of dynamic pages we index. In addition, our crawlers may suspect that a URL with many dynamic parameters might be the same page as another URL with different parameters. For that reason, we recommend using fewer parameters if possible. Typically, URLs with 1-2 parameters are more easily crawlable than those with many parameters. Also, you can help us find your dynamic URLs by submitting them to Google Sitemaps. To learn more, visit our FAQ at http://www.google.com/webmasters/sit...cs/en/faq.html.
You employ doorway pages. Google does not encourage the use of automatically generated pages that are designed for search engines instead of users. We want to point users to useful content pages, not to doorways or splash screens.
Your pages use frames. Google supports frames to the extent that we can. Frames tend to cause problems with search engines, bookmarks, emailing links and so on, because frames don't fit the conceptual model of the web (every page corresponds to a single URL). If a user's query matches the page as a whole, Google returns the frame set. If a user's query matches an individual frame on the page, Google returns the URL for that frame. The page is not displayed in a frame because there may be no frame set corresponding to that URL.
If you're concerned with the description of your site as seen by search engines, please read "Search Engines and Frames." It describes the "NoFrames" tag, which is used to provide alternative content. If, instead of providing alternative content, you use wording such as "This site requires the use of frames" or "Upgrade your browser," you're excluding both search engines and individuals whose browsers don't support frames. (For example, audio web browsers, such as those used in automobiles and by the visually impaired, typically do not deal with frames, which are a visual mechanism.) You can read about "NoFrames" in the HTML standard here.
3. Some of my pages are included, but others are missing.
Although we index billions of webpages, we cannot guarantee that we'll crawl all the pages on a particular site. However, we're always working to increase the number of pages we crawl and hope to include more pages in our index over time. For more information about how we find and include pages in our index please read our technology overview.
If your site's internal link structure doesn't provide a path to all of your pages, our robot may not see all the pages on your site. Google follows links from one page to the next, so pages that aren't linked to by others may be missed. Please see our Webmaster Guidelines for other ways to make your site more crawlable.
Although you can't buy your way into our search results, you can purchase advertising adjacent to them. More information about advertising with Google can be found here.
-
My site's listing is incorrect and I need it changed.
My site's listing is incorrect and I need it changed.
1. My information is outdated.
When you update information on your site, it doesn't instantly propagate to Google's index. Rather, Google's index is updated after our robots crawl a page. The crawl process is completely automated, so it's not necessary to submit updated or outdated links to us. Changes to your site's content will be noted when we next crawl your pages. Due to the volume of sites in our index, we cannot manually update pages on an individual basis.
2. I migrated my website to a new URL.
If you've changed your URL, or plan to, and would like Google to display your new URL, please keep in mind that we can't manually change your listed address in our search results. That said, there are steps you can take to make sure your transition is smooth.
If your old URLs redirect to your new site using HTTP 301 (permanent) redirects, our crawler will discover the new URLs. For more information about 301 HTTP redirects, please see http://www.ietf.org/rfc/rfc2616.txt.
Google listings are based in part on our ability to find you from links on other sites. To preserve your rank and help our crawler find your new URL, you'll want to inform others who link to you of your change of address. To find a sampling of sites that link to yours, perform a link search by entering "link:[your full URL]" into the Google search box. To find more pages that mention your URL, perform a Google search on your URL and select the "Find web pages that contain the term" link. Also, don't forget to change any entries you may have in directories such as Yahoo! or the Open Directory Project.
Finally, you may submit a list of your new URLs through the Google Sitemaps (Beta) program. Google Sitemaps uses webmaster-generated Sitemap files to learn about your webpages and to direct our crawlers to new and updated content.
Sometimes during site transitions, we'll fail to find a site at its new address. Just be sure that others are linking to you, and we should discover your new site.
3. There's no description of my site.
The Google index contains two types of pages -- fully indexed and partially indexed pages. Your page is currently partially indexed, which means that although we know about your site, our robots haven't read all the content on your pages in past crawls. This doesn't adversely affect your PageRank or your inclusion in our index. It does mean that we don't have detailed information about your page, so we display its URL as the title and omit a description. We understand the frustration this may cause you, and we're always working to increase the number of fully indexed pages in our search results.
4. The description of my site is wrong in the results.
Google's creation of snippets is completely automated and takes into account both the content of a page as well as references to it that appear on the web. We don't manually change sites' descriptions, but we're always working to make our snippets as relevant as possible.
-
I'm puzzled by my site's ranking.
1. How Google ranks pages.
Google's order of results is automatically determined by more than 100 factors, including our PageRank algorithm. Please check out our Technology Overview page for more details. Due to the nature of our business and our interest in protecting the integrity of our search results, we limit the information we make available to the public about our ranking system.
2. My page's location in the search results keeps changing.
Each time we update our database of webpages, our index invariably shifts: we find new sites, we lose some sites, and sites' ranking may change. Your rank naturally will be affected by changes in the ranking of other sites. No one at Google hand adjusts the results to boost the ranking of a site. The order of Google's search results is automatically determined by many factors, including our PageRank algorithm, and is described in more detail here.
You might check to see if the number of other sites that link to your URL has decreased. This is the single biggest factor in determining which sites are indexed by Google, as we find most pages when our robots crawl the web, jumping from page to page via hyperlinks. To find a sampling of sites that link to yours, try a Google link search.
3. My pages don't return for certain keywords.
Google does not manually assign keywords to sites, nor do we manually "boost" the rankings of any site. The ranking process is completely automated and takes into account more than 100 factors to determine the relevance of each result.
If you'd like your site to return for particular keywords, include these words on your pages. Our crawler analyzes the content of webpages in our index to determine the search queries for which they're most relevant. If your site clearly and accurately describes your topic and many other websites link to yours, it'll likely return as a search result for your desired keywords.
If you feel that certain keywords are essential to your site's success, you may want to consider our targeted keyword advertising program. Google does not sell placement in our results, but we do offer advertising adjacent to them. Please note that advertising with Google neither helps nor hurts your site's ranking in our search results
-
Webmaster Guidelines
Following these guidelines will help Google find, index, and rank your site. Even if you choose not to implement any of these suggestions, we strongly encourage you to pay very close attention to the "Quality Guidelines," which outline some of the illicit practices that may lead to a site being removed entirely from the Google index. Once a site has been removed, it will no longer show up in results on Google.com or on any of Google's partner sites.
Design and Content Guidelines:
Make a site with a clear hierarchy and text links. Every page should be reachable from at least one static text link.
Offer a site map to your users with links that point to the important parts of your site. If the site map is larger than 100 or so links, you may want to break the site map into separate pages.
Create a useful, information-rich site, and write pages that clearly and accurately describe your content.
Think about the words users would type to find your pages, and make sure that your site actually includes those words within it.
Try to use text instead of images to display important names, content, or links. The Google crawler doesn't recognize text contained in images.
Make sure that your TITLE and ALT tags are descriptive and accurate.
Check for broken links and correct HTML.
If you decide to use dynamic pages (i.e., the URL contains a "?" character), be aware that not every search engine spider crawls dynamic pages as well as static pages. It helps to keep the parameters short and the number of them few.
Keep the links on a given page to a reasonable number (fewer than 100).
-
Technical Guidelines:
Technical Guidelines:
Use a text browser such as Lynx to examine your site, because most search engine spiders see your site much as Lynx would. If fancy features such as JavaScript, cookies, session IDs, frames, DHTML, or Flash keep you from seeing all of your site in a text browser, then search engine spiders may have trouble crawling your site.
Allow search bots to crawl your sites without session IDs or arguments that track their path through the site. These techniques are useful for tracking individual user behavior, but the access pattern of bots is entirely different. Using these techniques may result in incomplete indexing of your site, as bots may not be able to eliminate URLs that look different but actually point to the same page.
Make sure your web server supports the If-Modified-Since HTTP header. This feature allows your web server to tell Google whether your content has changed since we last crawled your site. Supporting this feature saves you bandwidth and overhead.
Make use of the robots.txt file on your web server. This file tells crawlers which directories can or cannot be crawled. Make sure it's current for your site so that you don't accidentally block the Googlebot crawler. Visit http://www.robotstxt.org/wc/faq.html to learn how to instruct robots when they visit your site.
If your company buys a content management system, make sure that the system can export your content so that search engine spiders can crawl your site.
Don't use "&id=" as a parameter in your URLs, as we don't include these pages in our index
-
When your site is ready:
Have other relevant sites link to yours.
Submit it to Google at http://www.google.com/addurl.html.
Submit a sitemap as part of our Google Sitemaps (Beta) project. Google Sitemaps uses your sitemap to learn about the structure of your site and to increase our coverage of your webpages.
Make sure all the sites that should know about your pages are aware your site is online.
Submit your site to relevant directories such as the Open Directory Project and Yahoo!, as well as to other industry-specific expert sites.
-
Quality Guidelines - Basic principles:
Quality Guidelines - Basic principles:
Make pages for users, not for search engines. Don't deceive your users or present different content to search engines than you display to users, which is commonly referred to as "cloaking."
Avoid tricks intended to improve search engine rankings. A good rule of thumb is whether you'd feel comfortable explaining what you've done to a website that competes with you. Another useful test is to ask, "Does this help my users? Would I do this if search engines didn't exist?"
Don't participate in link schemes designed to increase your site's ranking or PageRank. In particular, avoid links to web spammers or "bad neighborhoods" on the web, as your own ranking may be affected adversely by those links.
Don't use unauthorized computer programs to submit pages, check rankings, etc. Such programs consume computing resources and violate our Terms of Service. Google does not recommend the use of products such as WebPosition Gold™ that send automatic or programmatic queries to Google.
Quality Guidelines - Specific recommendations:
Avoid hidden text or hidden links.
Don't employ cloaking or sneaky redirects.
Don't send automated queries to Google.
Don't load pages with irrelevant words.
Don't create multiple pages, subdomains, or domains with substantially duplicate content.
Avoid "doorway" pages created just for search engines, or other "cookie cutter" approaches such as affiliate programs with little or no original content.
These quality guidelines cover the most common forms of deceptive or manipulative behavior, but Google may respond negatively to other misleading practices not listed here (e.g. tricking users by registering misspellings of well-known websites). It's not safe to assume that just because a specific deceptive technique isn't included on this page, Google approves of it. Webmasters who spend their energies upholding the spirit of the basic principles listed above will provide a much better user experience and subsequently enjoy better ranking than those who spend their time looking for loopholes they can exploit.
If you believe that another site is abusing Google's quality guidelines, please report that site at http://www.google.com/contact/spamreport.html. Google prefers developing scalable and automated solutions to problems, so we attempt to minimize hand-to-hand spam fighting. The spam reports we receive are used to create scalable algorithms that recognize and block future spam attempts.
-
Remove your entire website
If you wish to exclude your entire website from Google's index, you can place a file at the root of your server called robots.txt. This is the standard protocol that most web crawlers observe for excluding a web server or directory from an index. More information on robots.txt is available here: http://www.robotstxt.org/wc/norobots.html. Please note that Googlebot does not interpret a 401/403 response ("Unauthorized"/"Forbidden") to a robots.txt fetch as a request not to crawl any pages on the site.
To remove your site from search engines and prevent all robots from crawling it in the future, place the following robots.txt file in your server root:
User-agent: *
Disallow: /
To remove your site from Google only and prevent just Googlebot from crawling your site in the future, place the following robots.txt file in your server root:
User-agent: Googlebot
Disallow: /
Each port must have its own robots.txt file. In particular, if you serve content via both http and https, you'll need a separate robots.txt file for each of these protocols. For example, to allow Googlebot to index all http pages but no https pages, you'd use the robots.txt files below.
For your http protocol (http://yourserver.com/robots.txt):
User-agent: *
Allow: /
For the https protocol (https://yourserver.com/robots.txt):
User-agent: *
Disallow: /
Note: If you believe your request is urgent and cannot wait until the next time Google crawls your site, use our automatic URL removal system. In order for this automated process to work, the webmaster must first create and place a robots.txt file on the site in question.
Google will continue to exclude your site or directories from successive crawls if the robots.txt file exists in the web server root. If you do not have access to the root level of your server, you may place a robots.txt file at the same level as the files you want to remove. Doing this and submitting via the automatic URL removal system will cause a temporary, 180 day removal of your site from the Google index, regardless of whether you remove the robots.txt file after processing your request. (Keeping the robots.txt file at the same level would require you to return to the URL removal system every 180 days to reissue the removal.)
Remove part of your website
Option 1: Robots.txt
To remove directories or individual pages of your website, you can place a robots.txt file at the root of your server. For information on how to create a robots.txt file, see the The Robot Exclusion Standard. When creating your robots.txt file, please keep the following in mind: When deciding which pages to crawl on a particular host, Googlebot will obey the first record in the robots.txt file with a User-agent starting with "Googlebot." If no such entry exists, it will obey the first entry with a User-agent of "*". Additionally, Google has introduced increased flexibility to the robots.txt file standard through the use asterisks. Disallow patterns may include "*" to match any sequence of characters, and patterns may end in "$" to indicate the end of a name.
To remove all pages under a particular directory (for example, lemurs), you'd use the following robots.txt entry:
User-agent: Googlebot
Disallow: /lemurs
To remove all files of a specific file type (for example, .gif), you'd use the following robots.txt entry:
User-agent: Googlebot
Disallow: /*.gif$
To remove dynamically generated pages, you'd use this robots.txt entry:
User-agent: Googlebot
Disallow: /*?
Option 2: Meta tags
Another standard, which can be more convenient for page-by-page use, involves adding a <META> tag to an HTML page to tell robots not to index the page. This standard is described at http://www.robotstxt.org/wc/exclusion.html#meta.
To prevent all robots from indexing a page on your site, you'd place the following meta tag into the <HEAD> section of your page:
<META NAME="ROBOTS" CONTENT="NOINDEX, NOFOLLOW">
To allow other robots to index the page on your site, preventing only Google's robots from indexing the page, you'd use the following tag:
<META NAME="GOOGLEBOT" CONTENT="NOINDEX, NOFOLLOW">
To allow robots to index the page on your site but instruct them not to follow outgoing links, you'd use the following tag:
<META NAME="ROBOTS" CONTENT="NOFOLLOW">
Note: If you believe your request is urgent and cannot wait until the next time Google crawls your site, use our automatic URL removal system. In order for this automated process to work, the webmaster must first insert the appropriate meta tags into the page's HTML code. Doing this and submitting via the automatic URL removal system will cause a temporary, 180-day removal of these pages from the Google index, regardless of whether you remove the robots.txt file or meta tags after processing your request.
Remove snippets
A snippet is a text excerpt that appears below a page's title in our search results and describes the content of the page..
To prevent Google from displaying snippets for your page, place this tag in the <HEAD> section of your page:
<META NAME="GOOGLEBOT" CONTENT="NOSNIPPET">
Note: removing snippets also removes cached pages.
Note: If you believe your request is urgent and cannot wait until the next time Google crawls your site, use our automatic URL removal system. In order for this automated process to work, the webmaster must first insert the appropriate meta tags into the page's HTML code.
Remove cached pages
Google automatically takes a "snapshot" of each page it crawls and archives it. This "cached" version allows a webpage to be retrieved for your end users if the original page is ever unavailable (due to temporary failure of the page's web server). The cached page appears to users exactly as it looked when Google last crawled it, and we display a message at the top of the page to indicate that it's a cached version. Users can access the cached version by choosing the "Cached" link on the search results page.
To prevent all search engines from showing a "Cached" link for your site, place this tag in the <HEAD> section of your page::
<META NAME="ROBOTS" CONTENT="NOARCHIVE">
To allow other search engines to show a "Cached" link, preventing only Google from displaying one, use the following tag:
<META NAME="GOOGLEBOT" CONTENT="NOARCHIVE">
Note: this tag only removes the "Cached" link for the page. Google will continue to index the page and display a snippet.
Note: If you believe your request is urgent and cannot wait until the next time Google crawls your site, use our automatic URL removal system. In order for this automated process to work, the webmaster must first insert the appropriate meta tags into the page's HTML code.
Remove an outdated ("dead") link
Google updates its entire index automatically on a regular basis. When we crawl the web, we find new pages, discard dead links, and update links automatically. Links that are outdated now will most likely "fade out" of our index during our next crawl.
Note: If you believe your request is urgent and cannot wait until the next time Google crawls your site, use our automatic URL removal system. We'll accept your removal request only if the page returns a true 404 error via the http headers. Please ensure that you return a true 404 error even if you choose to display a more user-friendly body of the HTML page for your visitors. It won't help to return a page that says "File Not Found" if the http headers still return a status code of 200, or normal.
Remove an image from Google's Image Search
To remove an image from Google's image index, add a robots.txt file to the root of the server. (If you can't put it in the server root, you can put it at directory level.)
Example: If you want Google to exclude the dogs.jpg image that appears on your site at www.yoursite.com/images/dogs.jpg, create a page at www.yoursite.com/robots.txt and add the following text:
User-agent: Googlebot-Image
Disallow: /images/dogs.jpg
To remove all the images on your site from our index, place the following robots.txt file in your server root:
User-agent: Googlebot-Image
Disallow: /
This is the standard protocol that most web crawlers observe for excluding a web server or directory from an index. More information on robots.txt is available here: http://www.robotstxt.org/wc/norobots.html.
Additionally, Google has introduced increased flexibility to the robots.txt file standard through the use asterisks. Disallow patterns may include "*" to match any sequence of characters, and patterns may end in "$" to indicate the end of a name. To remove all files of a specific file type (for example, to include .jpg but not .gif images), you'd use the following robots.txt entry:
User-agent: Googlebot-Image
Disallow: /*.gif$
Note: If you believe your request is urgent and cannot wait until the next time Google crawls your site, use our automatic URL removal system. In order for this automated process to work, the webmaster must first create and place a robots.txt file on the site in question.
Google will continue to exclude your site or directories from successive crawls if the robots.txt file exists in the web server root. If you do not have access to the root level of your server, you may place a robots.txt file at the same level as the files you want to remove. Doing this and submitting via the automatic URL removal system will cause a temporary, 180 day removal of the directories specified in your robots.txt file from the Google index, regardless of whether you remove the robots.txt file after processing your request. (Keeping the robots.txt file at the same level would require you to return to the URL removal system every 180 days to reissue the removal.)
Remove a blog from Blog Search
Only blogs with site feeds will be included in Blog Search. If you'd like to prevent your feed from being crawled, make use of a robots.txt file or meta tags (NOINDEX or NOFOLLOW), as described above. Please note that if you have a feed that was previously included, the old posts will remain in the index even though new ones will not be added.
Remove a RSS or Atom feed
When users add your feed to their Google homepage or Google Reader, Google's Feedfetcher attempts to obtain the content of the feed in order to display it. Since Feedfetcher requests come from explicit action by human users, Feedfetcher has been designed to ignore robots.txt guidelines.
It's not possible for Google to restrict access to a publicly available feed. If your feed is provided by a blog hosting service, you should work with them to restrict access to your feed. Check those sites' help content for more information (e.g., Blogger, LiveJournal, or Typepad).
Remove transcoded pages
Google Web Search on mobile phones allows users to search all the content in the Google index for desktop web browsers. Because this content isn't written specifically for mobile phones and devices and thus might not display properly, Google automatically translates (or "transcodes") these pages by analyzing the original HTML code and converting it to a mobile-ready format. To ensure that the highest quality and most useable web page is displayed on your mobile phone or device, Google may resize, adjust, or convert images, text formatting and/or certain aspects of web page functionality.
-
Nice work..
Thanks for Sharing this....
Posting Permissions
- You may not post new threads
- You may not post replies
- You may not post attachments
- You may not edit your posts
-
Forum Rules
Bookmarks