<DOC>
<DOCID>Navigating_the_Net.doc.txt</DOCID>
<AUTHOR>Doug Anger</AUTHOR>
<TEXT>
The Internet is a big place. The search engine Google had indexed over 8 billion pages at the last official count, and thats only a tiny fraction of whats really out there. (Google no longer displays a page count on their homepage.) Information in general is pretty easy to find, but finding the information you want can be like looking for a needle in a worldwide haystack.

Search Engines and Directories

The easiest way to search the net is with a search engine. Search engines like Google (http://www.google.com/) and Ask.com (http://search.ask.com/) create indices of web pages and allow you to search those indices for pages in several ways. These indices are created automatically by programs called spiders that crawl across the web, following every link they can find, and recording information about each page they encounter. In this manner, search engines can index billions of pages.

The Google search engine is currently the most popular on the web. It is a content-based search engine, meaning it searches the full text of every page in its index. Sometimes, however, it overlooks relevant pages. Google is famous primarily because of its PageRank algorithm that allows it to (theoretically) return the most relevant pages on the top of the results list based on the number of links to those pages from other pages on the Internet. Due to Googles popularity, however, an entire market has sprung up to manipulate the PageRank algorithm by changing the internal linking on a website or a group of websites. Google has been able to stop some of these schemes, but the PageRank algorithm (and therefore the search engine) remains partial to websites with certain linking patterns. In addition, Nevertheless, Google is without doubt one of the best search engines on the Internet.

Ask.com, while searching a smaller number of pages, has (in my opinion) a somewhat better method of searching pages for relevant results. Instead of using links to rank pages, its Teoma algorithm (also known as ExpertRank) uses links to sort them, providing the search engine what its makers call a 3D view of the web. Teoma makes intelligent suggestions for refining searches to give you results that are relevant to the topic you are researching. Although less popular than Google, Ask.com is also one of the best search engines on the Internet.

I would suggest starting with a basic search on Ask.com and using Teomas suggestions to refine your search. Then go to Google and use the new, refined search terms to locate any important pages that were not in Teomas index. It is never a bad idea to use two or more search engines.

Often confused with search engines are directories. A directory is a human-edited listing of web pages divided into topics. (Most directories are searchable, perhaps leading to the aforementioned confusion.) Directories generally contain fewer pages than search engines, often in the hundreds of thousands or low millions. Two popular directories are Yahoo! (http://www.yahoo.com/) and DMOZ (http://www.dmoz.org). If you are looking up a fairly common subject, a directory would be a good place to start because it would most likely contain a list of sites chosen by humans for that particular topic. However, if you are looking up a recent event, news item, person, or anything else that is not a common topic, a search engine would be your best bet. As with search engines, it is never a bad idea to look in multiple directories. It is a good idea to search in at least two search engines and two directories if you are doing any serious research on anything.

Specific Search Engines

If you wanted to find a pizza place in Sault Ste. Marie, would you go to the library and look it up in the card catalog? Of course not! So why would you use a general search engine to find specific information? If you are looking for something specific, there may be an easier way to find it. It is not possible to list all of the ways to find specific information, but Google has a few that are worth noting.

You can find information from college and universities using the Google University Search (http://www.google.com/options/universities.html), stand on the shoulders of giants by searching through scholarly papers with Google Scholar (http://scholar.google.com/), and find government pages from Googles Uncle Sam search (http://www.google.com/unclesam). If youre shopping, you can use Froogle (http://froogle.google.com/) to find the best prices and even search the mail order catalogs with Google Catalogs (http://catalogs.google.com/).

Databases

Web pages are a great source of information, but often, you will find even more by searching through databases. A database is a collection of records, each consisting of one or more fields. Records can be anything from archived newspaper articles to phone book entries.

A phone book is a very simple database. Each entry in the phone book is a record containing (usually) three fields: the name of the person or business listed, the address of that person or business, and the phone number of that person or business. In a phone book however, you can only search by name, not address, or phone number. You can look up the phone number and address associated with a particular name, but you cannot easily find the name and phone number of whoever lives at a particular address or the name and address of someone with a particular phone number. In an online database, you can usually search in any field or multiple fields at once. You could, for instance, find everyone named John who lived in Sault Ste. Marie and had a three in his phone number. (Try doing that with a phone book!) 

As a general rule, search engines and directories do not include database records in their results. For one reason, access to databases is often limited. For another, databases usually cover very specific subjects and contain many records on those subjects. It would not do to return half a million results from an aeronautic supply companys database of airplane parts every time someone looked up the phrase Boeing 747.

Michigan residents have access to the Michigan eLibrary (http://www.mel.org/). This virtual library consists of many databases which the state of Michigan has purchased access to. These include

The full text of over one hundred newspapers

About 11,440 full text books which can be read online (even more if you log in on campus)

Data on individuals from old US Census records

A worldwide library catalog

Thousands of other articles from magazines, newspapers, encyclopedias, and more



More databases are available through the LSSU library at http://www.lssu.edu/library/lib03/dblist.html.

Library of Congress

The Library of Congress (http://www.loc.gov/) provides free access to a number of databases and archives. The THOMAS system (http://thomas.loc.gov/) gives you access to bills that have come before congress in the current term by keyword or bill number. This should be a valuable resource to anyone studying current events.

The Library of Congress also provides digital archives of old documents and photographs and a number of other services. The American Memory collection alone contains over 7.5 million items.

Wiki

A wiki is a special type of website that anyone can edit. There have been a few devious people that have messed things up by adding inaccurate or biased information, but since most wikis allow anyone to undo any change instantly and let system administrators block the deviant from future access, these problems tend not to last long.  In the last few years, wikis have grown significantly.

Wikipedia (http://www.wikipedia.org/) is a wiki encyclopedia with over one million articles in several languages. Also available are the wiki dictionary know as Wiktionary (http://www.wiktionary.org/), and an online source of original nonfiction books called Wikibooks (http://www.wikibooks.org/), You can also find the latest news on Wikinews (http://www.wiknews.org/) or look up a species  animal, plant, or otherwise  with Wikispecies (http://species.wikipedia.org/).

Many professors do not accept wikis as legitimate sources, and for good reason. Since anyone can post to them, you dont know if your information is coming from someone with a Ph.D. or a prankster with a third-grade education. Wikis are an excellent resource to get you started on a particular subject, but it is always a good idea to go to the original sources cited in the article. 

Other Sources of Information

There are a few other sources of information worth noting that do not fit well into the categories we have discussed so far. Project Gutenberg (http://www.gutenberg.org/), named for the inventor of the printing press, is an archive of public domain books in electronic format. Gutenberg contains over 15,000 e-Books available for free download. Most of these are older books for which the copyright has expired.

The Internet Archive (http://www.archive.org/) contains a number of audio, video, and text documents for free access, but finding what youre looking for can be difficult. However, The Internet Archive also maintains an archive of websites as they appeared when they were indexed. Next time you encounter a 404 (file not found) error or a message informing you that the site no longer exists, copy its URL, paste it into the textbox at the Internet Archive, and see if the page did exist at one time.

Yahoo! offers a Creative Commons search (http://search.yahoo.com/cc) that allows you to find web pages, music, pictures, and movies that you can share freely, modify or build on, or even use commercially. Currently the Creative Commons search is dominated by weblogs, most of which are interesting, but have little or no educational or research value. There are, however a few good scientific and otherwise useful sites to be found and Creative Commons (http://www.creativecommons.org/) is trying to get encourage more people to put their educational works under creative commons licenses.



</TEXT></DOC>
