Followers

Serch For Your Favorites

Yahoo | Google | MSN | YouTube | Accoona | ASK | AlltheWeb | AltaVista | AllCrawl | Dewa | FindWhat | Search | YOUSEARCH < Dmoz | Lycos | IXquick | GO | HogSearch | HotBot | 7SEarch | AOL |

Saturday, 2 February 2008

-Types of Search Queries:

Andrei Broder authored A Taxonomy of Web Search [PDF], which notes that most searches fall into the following 3 categories:
Informational - seeking static information about a topic
Transactional - shopping at, downloading from, or otherwise interacting with the result
Navigational - send me to a specific URL
Improve Your Searching Skills:
Want to become a better searcher? Most large scale search engines offer:
Advanced search pages which help searchers refine their queries to request files which are newer or older, local or in nature, from specific domains, published in specific formats, or other ways of refining search, for example the ~ character means related to Google.
Vertical search databases which may help structure the information index or limit the search index to a more trusted or better structured collection of sources, documents, and information.
Nancy Blachman's Google Guide offers searchers free Google search tips, and Greg R.Notess's Search Engine Showdown offers a search engine features chart.
There are also many popular smaller vertical search services. For example, Del.icio.us allows you to search URLs that users have bookmarked, and Technorati allows you to search blogs.
World Wide Web Wanderer:
Soon the web's first robot came. In June 1993 Matthew Gray introduced the World Wide Web Wanderer. He initially wanted to measure the growth of the web and created this bot to count active web servers. He soon upgraded the bot to capture actual URL's. His database became knows as the Wandex.
The Wanderer was as much of a problem as it was a solution because it caused system lag by accessing the same page hundreds of times a day. It did not take long for him to fix this software, but people started to question the value of bots.

-Parts of a Search Engine:

Search engines consist of 3 main parts. Search engine spiders follow links on the web to request pages that are either not yet indexed or have been updated since they were last indexed. These pages are crawled and are added to the search engine index (also known as the catalog). When you search using a major search engine you are not actually searching the web, but are searching a slightly outdated index of content which roughly represents the content of the web. The third part of a search engine is the search interface and relevancy software. For each search query search engines typically do most or all of the following
Accept the user inputted query, checking to match any advanced syntax and checking to see if the query is misspelled to recommend more popular or correct spelling variations.
Check to see if the query is relevant to other vertical search databases (such as news search or product search) and place relevant links to a few items from that type of search query near the regular search results.
Gather a list of relevant pages for the organic search results. These results are ranked based on page content, usage data, and link citation data.
Request a list of relevant ads to place near the search results.
Searchers generally tend to click mostly on the top few search results, as noted in this article by Jakob Nielsen, and backed up by this search result eye tracking study.
Want to learn more about how search engines work?
In How does Google collect and rank results? Google engineer Matt Cutts briefly discusses how Google works.
Google engineer Jeff Dean lectures a University of Washington class on how a search query at Google works in this video.
The Chicago Tribune ran a special piece titled Gunning for Google, including around a dozen audio interviews, 3 columns, and this graphic about how Google works.
How Stuff Works covers search engines in How Internet Search Engines Work.

-History of Search Engines: From 1945 to Google 2007

As We May Think (1945):
The concept of hypertext and a memory extension really came to life in July of 1945, when after enjoying the scientific camaraderie that was a side effect of WWII, Vannaver Bush's As We May Think was published in The Atlantic Monthly.

He urged scientists to work together to help build a body of knowledge for all mankind. Here are a few selected sentences and paragraphs that drive his point home.
Specialization becomes increasingly necessary for progress, and the effort to bridge between disciplines is correspondingly superficial.
The difficulty seems to be, not so much that we publish unduly in view of the extent and variety of present day interests, but rather that publication has been extended far beyond our present ability to make real use of the record. The summation of human experience is being expanded at a prodigious rate, and the means we use for threading through the consequent maze to the momentarily important item is the same as was used in the days of square-rigged ships.
A record, if it is to be useful to science, must be continuously extended, it must be stored, and above all it must be consulted.
He not only was a firm believer in storing data, but he also believed that if the data source was to be useful to the human mind we should have it represent how the mind works to the best of our abilities.
Our ineptitude in getting at the record is largely caused by the artificiality of the systems of indexing. ... Having found one item, moreover, one has to emerge from the system and re-enter on a new path.
The human mind does not work this way. It operates by association. ... Man cannot hope fully to duplicate this mental process artificially, but he certainly ought to be able to learn from it. In minor ways he may even improve, for his records have relative permanency.
Presumably man's spirit should be elevated if he can better review his own shady past and analyze more completely and objectively his present problems. He has built a civilization so complex that he needs to mechanize his records more fully if he is to push his experiment to its logical conclusion and not merely become bogged down part way there by overtaxing his limited memory.
He then proposed the idea of a virtually limitless, fast, reliable, extensible, associative memory storage and retrieval system. He named this device a memex.
Gerard Salton (1960s - 1990s):
Gerard Salton, who died on August 28th of 1995, was the father of modern search technology. His teams at Harvard and Cornell developed the SMART informational retrieval system. Salton’s Magic Automatic Retriever of Text included important concepts like the vector space model, Inverse Document Frequency (IDF), Term Frequency (TF), term discrimination values, and relevancy feedback mechanisms.
He authored a 56 page book called A Theory of Indexing which does a great job explaining many of his tests upon which search is still largely based. Tom Evslin posted a blog entry about what it was like to work with Mr. Salton.
Ted Nelson:
Ted Nelson created Project Xanadu in 1960 and coined the term hypertext in 1963. His goal with Project Xanadu was to create a computer network with a simple user interface that solved many social problems like attribution.
While Ted was against complex markup code, broken links, and many other problems associated with traditional HTML on the WWW, much of the inspiration to create the WWW was drawn from Ted's work.
There is still conflict surrounding the exact reasons why Project Xanadu failed to take off.
The Wikipedia offers background and many resource links about Mr. Nelson.
Advanced Research Projects Agency Network:
ARPANet is the network which eventually led to the internet. The Wikipedia has a great background article on ARPANet and Google Video has a free interesting video about ARPANet from 1972.
Archie (1990):
The first few hundred web sites began in 1993 and most of them were at colleges, but long before most of them existed came Archie. The first search engine created was Archie, created in 1990 by Alan Emtage, a student at McGill University in Montreal. The original intent of the name was "archives," but it was shortened to Archie.
Archie helped solve this data scatter problem by combining a script-based data gatherer with a regular expression matcher for retrieving file names matching a user query. Essentially Archie became a database of web filenames which it would match with the users queries.
Bill Slawski has more background on Archie here.
Veronica & Jughead:
As word of mouth about Archie spread, it started to become word of computer and Archie had such popularity that the University of Nevada System Computing Services group developed Veronica. Veronica served the same purpose as Archie, but it worked on plain text files. Soon another user interface name Jughead appeared with the same purpose as Veronica, both of these were used for files sent via Gopher, which was created as an Archie alternative by Mark McCahill at the University of Minnesota in 1991.
File Transfer Protocol:
Tim Burners-Lee existed at this point, however there was no World Wide Web. The main way people shared data back then was via File Transfer Protocol (FTP).
If you had a file you wanted to share you would set up an FTP server. If someone was interested in retrieving the data they could using an FTP client. This process worked effectively in small groups, but the data became as much fragmented as it was collected.
Tim Berners-Lee & the WWW (1991):
From the Wikipedia:
While an independent contractor at CERN from June to December 1980, Berners-Lee proposed a project based on the concept of hypertext, to facilitate sharing and updating information among researchers. With help from Robert Cailliau he built a prototype system named Enquire.
After leaving CERN in 1980 to work at John Poole's Image Computer Systems Ltd., he returned in 1984 as a fellow. In 1989, CERN was the largest Internet node in Europe, and Berners-Lee saw an opportunity to join hypertext with the Internet. In his words, "I just had to take the hypertext idea and connect it to the TCP and DNS ideas and — ta-da! — the World Wide Web". He used similar ideas to those underlying the Enquire system to create the World Wide Web, for which he designed and built the first web browser and editor (called WorldWideWeb and developed on NeXTSTEP) and the first Web server called httpd (short for HyperText Transfer Protocol daemon).
The first Web site built was at http://info.cern.ch/ and was first put online on August 6, 1991. It provided an explanation about what the World Wide Web was, how one could own a browser and how to set up a Web server. It was also the world's first Web directory, since Berners-Lee maintained a list of other Web sites apart from his own.
In 1994, Berners-Lee founded the World Wide Web Consortium (W3C) at the Massachusetts Institute of Technology.
Tim also created the Virtual Library, which is the oldest catalogue of the web. Tim also wrote a book about creating the web, titled Weaving the Web.
What is a Bot?
Computer robots are simply programs that automate repetitive tasks at speeds impossible for humans to reproduce. The term bot on the internet is usually used to describe anything that interfaces with the user or that collects data.
Search engines use "spiders" which search (or spider) the web for information. They are software programs which request pages much like regular browsers do. In addition to reading the contents of pages for indexing spiders also record links.
Link citations can be used as a proxy for editorial trust.
Link anchor text may help describe what a page is about.
Link co citation data may be used to help determine what topical communities a page or website exist in.
Additionally links are stored to help search engines discover new documents to later crawl.
Another bot example could be Chatterbots, which are resource heavy on a specific topic. These bots attempt to act like a human and communicate with humans on said topic.

-Multiple search engines

Multiple search engines allow you to use many engines at once. The good news is that this is a fast and comprehensive way to cast a wide net. The bad news is that you cannot always narrow your search effectively because different engines use different codes. So use these multiple searches if you are hunting for something obscure and cannot find enough hits when you use your favourite, single search engine.

-Comparison of the search engines.

I am not going to pretend that this comparison is the result of highly scientific work; I have also not checked my results with the producers of the different engines, though I have tried to ensure accuracy. Due to the fast changing nature of the Internet however, you are advised to check the information for yourself, since it will doubtless become out of date quickly.I chose Meta-search engines which offered as much variety as possible, and gave me examples of list, consecutive and simultaneous searches, and obtained the data for all of them on 25th February 2003. According to Yahoo there are at present a total of over 100 different sites offering engines of this type, so for a full list I would point you to them.My conclusions, which are my own personal impressions, nothing more, are as follows.I see no value whatsoever in the list approach. These are not examples of what I would regard as 'true' multi-search engines; anyone could put a list of these together and claim that they had created a multi-search engine when in actual fact all that they have done is be slightly creative with cut and paste facilities. The possible exception is Metasearch, which does automatically put your desired search terms into the appropriate places on the cut and paste search engines they have referenced. This approach also is unable to properly provide boolean operators etc, since the page does not interact with the search engines themselves, simply providing this front end cut and paste job. Worse, they are usually unable to offer much by the way of help screens, since this is dependant on the search engines themselves.
Consecutive multi-search engines were however much better. They did make attempts to integrate their page into the search engines, and so are generally better at providing a wider range of functions, although I still found that help screens and guides to searching were very limited. The major disadvantage of this approach is that it can take considerable time for the search to be completed, and the weak link is always going to be the slowest engine that they reference. However, they do seem to work reasonably well, and are certainly worth experimenting with.
Simultaneous search engines seem to be few and far between, but they are without a doubt the most effective. Superseek uses the Frames approach to overcome the problem of obtaining and displaying results on the screen, but this approach does mean that you have to have a frames compatible browser available, which not everyone will have. The search results screen also looks as though its come straight out of an aeroplane cockpit and is a little daunting when you first view it. However, it does not take long to get used to. They do also have a non-frames approach, but I did not try this out. Worth experimenting with.
My two favourites however are the Internet Sleuth and Savvy Search. Both were helpful, fast and efficient. I would be quite happy to use either or both of these to run a multi-search, and I would recommend them.
I would welcome comments, additions, updates and so on; please feel free to email me.

-Characteristics of Multi-search engines.

I have tried to put together a list of the different elements which one might expect to appear on a multi-search engine page. Unfortunately few, if any of the multi-search engines exhibit all of these elements, and indeed some will have very few of them.
The number of search engines that a Multi-search engine will use
The number of search engines which are used varies dramatically - the smallest in the sample that I looked at only referred to half a dozen, while the largest in the sample gave me access to over one thousand search engines, or database front ends. This is no real indicator of quality however; it depends much more on the search engines which are used (and also the variety) rather than the sheer number. I would much prefer to use a Multi-search engine which referenced a small number of what I would consider to be high quality search engines than a much larger number of engines that I did not really know or trust.
Nonetheless, I think it is an acceptable criteria to use when evaluating the effectiveness of a Multi-search engine; while more is not necessarily better, less could certainly be considered worse. While I would not normally evaluate success in terms of the number of hits this is one of the reasons why one would use a Multi-search engine, so I feel that it is justified.
The elements of the Internet which are searched.
It seems almost automatic these days to regard the phrase 'Search the Web' as a synonym for 'Search the Internet'. Of course, while that is understandable, given the hype and attention surrounding the World Wide Web, it behoves us to remember that there are a number of other aspects of the Internet which deserve consideration as well, such as Usenet newsgroups, or individuals email addresses and so on. Multi-search engines are at the mercy of the search engines they choose to reference, but given that there are a good number of these which concentrate on specific aspects such as those just mentioned there is no reason why they should not be made available for searching as well.
Any words, all words, phrase searching.
Again, there is little that the Multi-search engine can do directly about this since they are unable to affect the internal workings of individual search engines. However, it is an option that should be offered to the end user; if one search engine can search on a phrase out of the list available it seems short sighted not to offer this. Other engines on the list will simply ignore the phrase aspect and search on the words using an implied OR. If this is not given as an option though, it reduces the effectiveness of those search engines which can undertake phrase searching. It seems to be an obvious point, but is one which was overlooked by some of the Multi-search engines which I looked at
Boolean operators, truncation and proximity searching.
The very same comments can be made here as just given above. A Multi-search engine should provide and reflect the variety of approaches made available by the engines it references, but all too often this was not the case.
Focussing a search.
There are of course many occasions when the user will not wish to do a global search, but will want to focus on one aspect specifically, such as searching in a specific domain (such as .com) or a geographical area (such as Europe). Yet again some search engines allow the user to focus the search this way, but this is not always reflected in the interface offered by the Multi-search engine.
Choice of subject area.
This approach is very familiar to anyone who has ever used the Index approach to search engines, by taking a broad subject area and choosing various subheadings until the specific subject is reached. It is well known to all of us in the information profession that much of the time we will not want everything on a subject, but will wish to focus on the medical or legal aspects for example. Some Multi-search engines did offer the ability to choose a specific subject focus or indeed provide a subset of search engines which cover a particular subject area, and the The Big Hub deserves to be singled out here for being quite superb in this approach.
Time taken and hits returned.
Two important elements here which both approach the problem of the amount of time that it takes to run some of these searches. A major disadvantage of using a Multi-search engine is that you are very much left in the hands of the engines referenced. If they decide to take a long time to return a result, or they are particularly comprehensive you may be sat twiddling your thumbs while they work. By limiting the search either by number of hits or by getting the search to end after a particular period of time you are able to exert at least a little control over the whole process. The danger here of course is that you are not going to get the same level of comprehensiveness that you might get otherwise, but at least you are being given the choice
Display
Many search engines will allow you one of three display modes; brief, normal or verbose. Your choice is likely going to depend on what you actually want from the search, and the variety can be quite helpful in some circumstances.
Collate results.
When using a single search engine it is not uncommon to find the same site turning up as a hit several times - de-duping does not seem to be that much of a priority for a lot of developers! This problem is exacerbated when you do a multi-search; it can be annoying to retrieve what appears to be a reasonable number of hits, only to find that there are a great many duplicates in the list. In my opinion, one of the keys strengths of a good Multi-search engine should be that it is able to collate the results, de-dupe and then display. Unfortunately however, it appeared to be very uncommon to find one which did.
Help screens/FAQ's
It is slightly distressing to see so many Internet retrieval engines attempt to give users the impression that searching is a very simple process, when of course we know that its rather more complex than that. All too often I would log onto a Multi-search engine page to find no instructions, no hints on how to create a better or more effective search, and no way of identifying how many search engines were used. Indeed, the lack of such information was quite astonishing. Perhaps its just me, but if I'd taken a lot of time to create a Multi-search interface, I'd want everyone to know who I was, how I had done it and why it was effective. Its possible that the producers of these facilities have rather smaller egos than I, but I tend to think that it is related rather more to laziness than anything else.

-How do Multi-search engines work?

From the explorations that I have undertaken, there appear to be 3 different approaches which are in operation at the moment:
A straightforward list of different search engines
Searches which take place one after another.
Searches which take place simultaneously
Each of these has their own advantages and disadvantages, so lets examine them in a little more detail.
A straightforward list of different search engines
These work by simply copying the appropriate URL for the cgi script onto the web page. This is not particularly difficult to do and the user then simply inputs the appropriate search term(s) into the dialogue box and submits the search. This is then run by the search engine in particular and the user is presented with a list of results in exactly the same way that they would if they had gone to visit that particular site directly.
The advantage of this approach is simply that you can reduce the amount of time you spend going from one site to another in order to complete your search.
It might also suggest other search engines for you which you had not considered using before. It is almost impossible to keep up with all of the different search engines which are available, and if someone is happy to do this on your behalf, it makes sense to take advantage of it.
However, the disadvantage of this approach is that, strictly speaking, these sites are (in my opinion) misleading the user. They are not offering a multi-search engine, but have simply collated the work of others onto a new home page. Thats not to say there is anything intrinsically wrong with this, and the Webmaster will have had to have done a reasonable amount of work to set the page up, but its really nothing more than a slightly more sophisticated list of links.
These are by far the most common sites offering a multi-search facility and examples of sites which take this approach are:
Find-It at http://www.itools.com/search/
Searches which take place one after another.
This is much closer to the concept of a multi-search engine. A site of this nature will usually have a single entry line where you input the search just as you would with a single search engine interface. You may then have the opportunity of deciding which search engines you the search to run under (usually from a list given in a check-box type situation) and the multi-search engine then transmits the search simultaneously to all of the search engines you have indicated.
Once the search has been run, the results will be displayed on the screen for you in a list, commonly separated into the results as provided by the different search engines. However, the main disadvantage of this type of approach is that before the list can be generated on the screen all the different search engines have to have sent their results back to the multi-search engine site. Consequently the speed of the search is dictated by the speed of the slowest search engine.
An example of this type of search engine can be found at:
Dogpile at http://www.dogpile.com/
Searches which take place simultaneously
This type of search engine is very similiar to the previous approach, the main difference being that searches do not have to wait until each search engine has completed its work - as soon as results are available from one search engine, they are displayed on the screen for you to view. The result is therefore that much faster, and while you are browsing down through the list of hits, others are being added to the page even as you view.
An example of the multi-search engines which take this approach are:
Ixquick at http://www.ixquick.com/