From ALA TechSource - How OPACs Suck, Part 1
http://www.techsource.ala.org/blog/2006/03/how-opacs-suck-part-1-relevance-rank-or-the-lack-of-it.html
So How Do You Make Cream, Anyway?
Relevance ranking is actually fairly simple technology. It's primarily determined by the magic of something every search-engine vendor will talk your ears off about: TF/IDF.
TF, for term frequency, measures the importance of the term in the item you're retrieving, whether you're searching a full-text book in Google or a catalog record. The more the term million shows up in the document—think of a catalog record for the book Million Little Pieces—the more important the term million is to the document.
IDF, for inverse document frequency, measures the importance of the word in the database you're searching. The fewer times the term million shows up in the entire database, the more important, or unique, it is.
Put TF and IDF together—the importance of a term in a document, and the uniqueness of the same term in an entire database—and you have basic relevance ranking. If the word million shows up several times in a catalog record, and it's not that common in the database, the item should rise to the top, as Endeca presents them in the NCSU catalog.
No comments:
Post a Comment