Topic 4 W4 - Text Processing

Uploaded by

VISALINI VIJAYAN

Available Formats

Download as PPTX, PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

11 views

Topic 4 W4 - Text Processing

Uploaded by

VISALINI VIJAYAN

Available Formats

Download as PPTX, PDF, TXT or read online on Scribd

You are on page 1/ 42

Search Engines

Information Retrieval in Practice

All slides ©Addison Wesley, 2008

Processing Text
• Converting documents to index terms
• Why?
– Matching the exact string of characters typed by
the user is too restrictive
• i.e., it doesn’t work very well in terms of effectiveness
– Not all words are of equal value in a search
– Sometimes not clear where words begin and end
• Not even clear what a word is in some languages
– e.g., Chinese, Korean
Text Statistics
• Huge variety of words used in text but
• Many statistical characteristics of word
occurrences are predictable
– e.g., distribution of word counts
• Retrieval models and ranking algorithms
depend heavily on statistical properties of
words
– e.g., important words occur often in documents
but are not high frequency in collection
Zipf’s Law
• Distribution of word frequencies is very
skewed
– a few words occur very often, many words hardly
ever occur
– e.g., two most common words (“the”, “of”) make
up about 10% of all word occurrences in text
documents
Zipf’s Law
• the frequency of any word is inversely proportional to its
rank in the frequency table.
• most frequent word will occur approximately twice as often
as the second most frequent word,
• three times as often as the third most frequent word, etc.
• For example in a doc., the word "the" is the most frequently
occurring word, and by itself accounts for nearly 7% of all
word occurrences (69,971 out of slightly over 1 million).
• True to Zipf's Law, the second-place word "of" accounts for
slightly over 3.5% of words (36,411 occurrences), followed
by "and" (28,852).
Zipf’s Law
Vocabulary Growth
• As corpus grows, so does vocabulary size
– Fewer new words when corpus is already large
• Observed relationship (Heaps’ Law):

v = k.nβ
where v is vocabulary size (number of unique words),
n is the number of words in corpus, k, β are
parameters that vary for each corpus (typical values
given are 10 ≤ k ≤ 100 and β ≈ 0.5)
AP89 Example

k β
Heaps’ Law Predictions
• Predictions for TREC collections are accurate
for large numbers of words
– e.g., first 10,879,522 words of the AP89 collection
scanned
– prediction is 100,151 unique words
– actual number is 100,024
• Predictions for small numbers of words (i.e.
< 1000) are much worse
GOV2 (Web) Example
Web Example
• Heaps’ Law works with very large corpora
– new words occurring even after seeing 30 million!
– parameter values different than typical TREC
values
• New words come from a variety of sources
• spelling errors, invented words (e.g. product, company
names), code, other languages, email addresses, etc.
• Search engines must deal with these large and
growing vocabularies
Estimating Result Set Size

• How many pages contain all of the query terms?

• For the query “a b c”:
fabc = N · fa/N · fb/N · fc/N = (fa · fb · fc)/N2

• Assuming that terms occur independently

• fabc is the estimated size of the result set
• fa, fb, fc are the number of documents that terms a, b, and
c occur in
• N is the number of documents in the collection
GOV2 Example

(fa · fb )/N

Collection size (N) is 25,205,179

Tokenizing
• Forming words from sequence of characters
• Surprisingly complex in English, can be harder
in other languages
• Early IR systems:
– any sequence of alphanumeric characters of
length 3 or more
– terminated by a space or other special character
– upper-case changed to lower-case
Tokenizing
• Example:
– “Bigcorp's 2007 bi-annual report showed profits
rose 10%.” becomes
– “bigcorp 2007 annual report showed profits rose”
• Why? Too much information lost
– Small decisions in tokenizing can have major
impact on effectiveness of some queries
Tokenizing Problems
• Small words can be important in some queries,
usually in combinations
• xp, ma, pm, ben e king, el paso, master p, gm, j lo, world
war II
• Both hyphenated and non-hyphenated forms of
many words are common
– Sometimes hyphen is not needed
• e-bay, wal-mart, active-x, cd-rom, t-shirts
– At other times, hyphens should be considered either
as part of the word or a word separator
• winston-salem, mazda rx-7, e-cards, pre-diabetes, t-mobile,
spanish-speaking
Tokenizing Problems
• Special characters are an important part of tags,
URLs, code in documents
• Capitalized words can have different meaning
from lower case words
– Bush, Apple
• Apostrophes can be a part of a word, a part of a
possessive, or just a mistake
– rosie o'donnell, can't, don't, 80's, 1890's, men's straw
hats, master's degree, england's ten largest cities,
shriner's
Tokenizing Problems
• Numbers can be important, including decimals
– nokia 3250, top 10 courses, united 93, quicktime
6.5 pro, 92.3 the beat, 288358
• Periods can occur in numbers, abbreviations,
URLs, ends of sentences, and other situations
– I.B.M., Ph.D., cs.umass.edu, F.E.A.R.
• Note: tokenizing steps for queries must be
identical to steps for documents
Tokenizing Process
• First step is to use parser to identify
appropriate parts of document to tokenize
• Defer complex decisions to other components
– word is any sequence of alphanumeric characters,
terminated by a space or special character, with
everything converted to lower-case
– everything indexed
– example: 92.3 → 92 3 but search finds documents
with 92 and 3 adjacent
Tokenizing Process
• Not that different than simple tokenizing
process used in past
• Examples of rules used with TREC
– Apostrophes in words ignored
• o’connor → oconnor bob’s → bobs
– Periods in abbreviations ignored
• I.B.M. → ibm Ph.D. → ph d
Stopping
• Function words (determiners, prepositions)
have little meaning on their own
• High occurrence frequencies
• Treated as stopwords (i.e. removed)
– reduce index space, improve response time,
improve effectiveness
• Can be important in combinations
– e.g., “to be or not to be”
Stopping
• Stopword list can be created from high-
frequency words or based on a standard list
• Lists are customized for applications, domains,
and even parts of documents
– e.g., “click” is a good stopword for anchor text
• Best policy is to index all words in documents,
make decisions about which words to use at
query time
Stemming
• Many morphological variations of words
– inflectional (plurals, tenses)
– derivational (making verbs nouns etc.)
• In most cases, these have the same or very
similar meanings
• Stemmers attempt to reduce morphological
variations of words to a common stem
– usually involves removing suffixes
• Can be done at indexing time or as part of
query processing (like stopwords)
Stemming
• Generally a small but significant effectiveness
improvement
– can be crucial for some languages
– e.g., 5-10% improvement for English, up to 50% in
Arabic

Words with the Arabic root ktb

Stemming
• Two basic types
– Dictionary-based: uses lists of related words
– Algorithmic: uses program to determine related
words
• Algorithmic stemmers
– suffix-s: remove ‘s’ endings assuming plural
• e.g., cats → cat, lakes → lake, wiis → wii
• Many false negatives: supplies → supplie
• Some false positives: ups → up
Porter Stemmer
• Algorithmic stemmer used in IR experiments
since the 70s
• Consists of a series of rules designed to the
longest possible suffix at each step
• Effective in TREC
• Produces stems not words
• Makes a number of errors and difficult to
modify
Krovetz Stemmer
• Hybrid algorithmic-dictionary
– Word checked in dictionary
• If present, either left alone or replaced with “exception”
• If not present, word is checked for suffixes that could be
removed
• After removal, dictionary is checked again
• Produces words not stems
• Comparable effectiveness
• Lower false positive rate, somewhat higher false
negative
Stemmer Comparison
Phrases
• Many queries are 2-3 word phrases
• Phrases are
– More precise than single words
• e.g., documents containing “black sea” vs. two words
“black” and “sea”
– Less ambiguous
• e.g., “big apple” vs. “apple”
• Can be difficult for ranking
• e.g., Given query “fishing supplies”, how do we score
documents with
– exact phrase many times, exact phrase just once, individual words
in same sentence, same paragraph, whole document, variations on
words?
Document Structure and Markup
• Some parts of documents are more important
than others
• Document parser recognizes structure using
markup, such as HTML tags
– Headers, anchor text, bolded text all likely to be
important
– Metadata can also be important
– Links used for link analysis
Example Web Page
hypertext
Example Web Page

hypertext
Link Analysis
• Links are a key component of the Web
• Important for navigation, but also for search
– e.g., <a href="http://example.com" >Example
website</a>
– “Example website” is the anchor text
– “http://example.com” is the destination link
– both are used by search engines
Anchor Text
• Used as a description of the content of the
destination page
– i.e., collection of anchor text in all links pointing to
a page used as an additional text field
• Anchor text tends to be short, descriptive, and
similar to query text
• Retrieval experiments have shown that anchor
text has significant impact on effectiveness for
some types of queries
PageRank
• Billions of web pages, some more informative
than others
• Links can be viewed as information about the
popularity (authority?) of a web page
– can be used by ranking algorithm
• Inlink count could be used as simple measure
• Link analysis algorithms like PageRank provide
more reliable ratings
Dangling Links
• Random jump prevents getting stuck on
pages that
– do not have links
– contains only links that no longer point to
other pages
– have links forming a loop
• Links that point to the first two types of
pages are called dangling links
– may also be links to pages that have not yet
been crawled
Link Quality
• Link quality is affected by spam and other
factors
– e.g., link farms to increase PageRank
– trackback links in blogs can create loops
– links from comments section of popular blogs
• Blog services modify comment links to contain
rel=nofollow attribute
• e.g., “Come visit my <a rel=nofollow
href="http://www.page.com">web page</a>.”
Trackback Links
Internationalization
• 2/3 of the Web is in English
• About 50% of Web users do not use English as
their primary language
• Many (maybe most) search applications have
to deal with multiple languages
– monolingual search: search in one language, but
with many possible languages
– cross-language search: search in multiple
languages at the same time
Internationalization
• Many aspects of search engines are language-
neutral
• Major differences:
– Text encoding (converting to Unicode)
– Tokenizing (many languages have no word
separators)
– Stemming
• Cultural differences may also impact interface
design and features provided
Chinese “Tokenizing”
END

Thimble - 2.1.3.zip - TXT Filename UTF-8''thimble 2.1.3.zip
No ratings yet
Thimble - 2.1.3.zip - TXT Filename UTF-8''thimble 2.1.3.zip
12 pages
Chap 4
No ratings yet
Chap 4
76 pages
1. 2_text Operation_1 (2)
No ratings yet
1. 2_text Operation_1 (2)
28 pages
Lecture 1 Text Preprocessing PDF
No ratings yet
Lecture 1 Text Preprocessing PDF
29 pages
AI6122 Topic 1.2 - WordLevel
No ratings yet
AI6122 Topic 1.2 - WordLevel
63 pages
Chapter 2 Part 1 & 2
No ratings yet
Chapter 2 Part 1 & 2
58 pages
Lecture 3
No ratings yet
Lecture 3
70 pages
Chapter 2 (Information Storage & Retrieval)
No ratings yet
Chapter 2 (Information Storage & Retrieval)
56 pages
Module5-Representing and Mining Text
No ratings yet
Module5-Representing and Mining Text
24 pages
CH 2_text operation
No ratings yet
CH 2_text operation
38 pages
Information Retrieval: Text Processing
No ratings yet
Information Retrieval: Text Processing
43 pages
NLP - 1_250119_222702 (1)
No ratings yet
NLP - 1_250119_222702 (1)
71 pages
6 The Term Vocabulary & Posting List
No ratings yet
6 The Term Vocabulary & Posting List
19 pages
IRS Chapter 2
No ratings yet
IRS Chapter 2
57 pages
DSB - Unit4-Representing and Miniing text-decision-analytic-think-II
No ratings yet
DSB - Unit4-Representing and Miniing text-decision-analytic-think-II
46 pages
NLP Lect-6 03.02.21
No ratings yet
NLP Lect-6 03.02.21
17 pages
AP for NLP-Word 2 Vec
No ratings yet
AP for NLP-Word 2 Vec
33 pages
NLP Lect-5 02.02.21
No ratings yet
NLP Lect-5 02.02.21
18 pages
10 - POS Tagging
No ratings yet
10 - POS Tagging
75 pages
Week3
No ratings yet
Week3
15 pages
Natural Language Processing
No ratings yet
Natural Language Processing
17 pages
Introduction To NLP
No ratings yet
Introduction To NLP
68 pages
Extracting, Cleaning and Pre-Processing Text
No ratings yet
Extracting, Cleaning and Pre-Processing Text
12 pages
AP for NLP-LO1
No ratings yet
AP for NLP-LO1
61 pages
NLP
No ratings yet
NLP
17 pages
IR Chapter 2 Text Operations
No ratings yet
IR Chapter 2 Text Operations
25 pages
2_text operation
No ratings yet
2_text operation
35 pages
CCS369 - TSS-Unit 4
No ratings yet
CCS369 - TSS-Unit 4
30 pages
Lec 19
No ratings yet
Lec 19
60 pages
2 TextOperations
No ratings yet
2 TextOperations
54 pages
6. Applications of NLP
No ratings yet
6. Applications of NLP
85 pages
Unit 1b
No ratings yet
Unit 1b
24 pages
Natural Language Processing
No ratings yet
Natural Language Processing
27 pages
NLP_Week_02
No ratings yet
NLP_Week_02
55 pages
Text Preprocessing
No ratings yet
Text Preprocessing
59 pages
NLP 3-6
No ratings yet
NLP 3-6
20 pages
Chapter 4
No ratings yet
Chapter 4
72 pages
NLP SEM QUESTIONS AND ANSWERS
No ratings yet
NLP SEM QUESTIONS AND ANSWERS
72 pages
chapter two IR
No ratings yet
chapter two IR
44 pages
Word Segmentation Sentence Segmentation: Recommended Reading
No ratings yet
Word Segmentation Sentence Segmentation: Recommended Reading
31 pages
Business Intelligence and Data Mining: by Dr. Atanu Rakshit Email: Atanu - Rakshit@iimrohtak - Ac.in
No ratings yet
Business Intelligence and Data Mining: by Dr. Atanu Rakshit Email: Atanu - Rakshit@iimrohtak - Ac.in
122 pages
2-Text Operations_new
No ratings yet
2-Text Operations_new
39 pages
feature eng
No ratings yet
feature eng
34 pages
NLP UNIT 5 part b
100% (2)
NLP UNIT 5 part b
31 pages
Kuhlmann - Introduction To Computational Linguistics (Slides) (2015)
100% (1)
Kuhlmann - Introduction To Computational Linguistics (Slides) (2015)
66 pages
NLP Part1
No ratings yet
NLP Part1
67 pages
IR Chap7
No ratings yet
IR Chap7
30 pages
9-Word and Sentence Segmentation-17!01!2024
No ratings yet
9-Word and Sentence Segmentation-17!01!2024
32 pages
Natural Language Processing
No ratings yet
Natural Language Processing
72 pages
Chap - Week8 - Queries and Information Needs
No ratings yet
Chap - Week8 - Queries and Information Needs
44 pages
Tidy Text
No ratings yet
Tidy Text
39 pages
Unit II
No ratings yet
Unit II
61 pages
NLP Unit-Ii
No ratings yet
NLP Unit-Ii
118 pages
Natural Language Processing 1
No ratings yet
Natural Language Processing 1
19 pages
NLP_Lecture_6_Week_3
No ratings yet
NLP_Lecture_6_Week_3
9 pages
DLT Unit-5
No ratings yet
DLT Unit-5
48 pages
NLP-Lectures 4,5,6
No ratings yet
NLP-Lectures 4,5,6
85 pages
NLP UNIT-II
No ratings yet
NLP UNIT-II
71 pages
3. Syntax Parsing
No ratings yet
3. Syntax Parsing
95 pages
Schematron: A language for validating XML
From Everand
Schematron: A language for validating XML
Erik Siegel
No ratings yet
Key & Common Swedish Words A Vocabulary List of High Frequency Swedish Words(1000 Words): Swedish, #0
From Everand
Key & Common Swedish Words A Vocabulary List of High Frequency Swedish Words(1000 Words): Swedish, #0
MostUsedWords
2/5 (4)
FM1100 Simple User Guide For Recommended Configuration V2.0
No ratings yet
FM1100 Simple User Guide For Recommended Configuration V2.0
8 pages
Zak - Ch03 Theory
No ratings yet
Zak - Ch03 Theory
51 pages
4.2.1 STS 3113 202020 Course Project
No ratings yet
4.2.1 STS 3113 202020 Course Project
5 pages
InfoSciV26p039 068morandini88951
No ratings yet
InfoSciV26p039 068morandini88951
31 pages
TL1 Reference
No ratings yet
TL1 Reference
456 pages
DOS - Presentation
No ratings yet
DOS - Presentation
44 pages
Distributed Control System & Scada: Chapter-3
No ratings yet
Distributed Control System & Scada: Chapter-3
33 pages
Innodisk m5s0 Bgm2oavp-3317554
No ratings yet
Innodisk m5s0 Bgm2oavp-3317554
22 pages
Scrum Master Resume Sample - Windsor Original
No ratings yet
Scrum Master Resume Sample - Windsor Original
1 page
HCPL-2630 FairchildSemiconductor PDF
No ratings yet
HCPL-2630 FairchildSemiconductor PDF
11 pages
4.4.4 Lab Locating Log Files
No ratings yet
4.4.4 Lab Locating Log Files
17 pages
Funny Pic - Google Search
No ratings yet
Funny Pic - Google Search
1 page
Maxtar 200 La-Zz
No ratings yet
Maxtar 200 La-Zz
142 pages
Os Lab Manual 7 -9 Programs
No ratings yet
Os Lab Manual 7 -9 Programs
17 pages
Lab Chapter # 5
No ratings yet
Lab Chapter # 5
7 pages
An Introduction To Basics of Interfaces in Oracle Apps
No ratings yet
An Introduction To Basics of Interfaces in Oracle Apps
15 pages
AUS Health Sample
No ratings yet
AUS Health Sample
35 pages
APPLICATIN DRAFT
No ratings yet
APPLICATIN DRAFT
6 pages
Interactive Mobile Technologies
No ratings yet
Interactive Mobile Technologies
12 pages
HT-150600 - PS Elite 750 Info
No ratings yet
HT-150600 - PS Elite 750 Info
4 pages
E-Note SS Two 3rd Term Data Processing
No ratings yet
E-Note SS Two 3rd Term Data Processing
14 pages
AssocDev Activities March2020
100% (1)
AssocDev Activities March2020
140 pages
PT10 20 - Mobile - Pentesting - Preview
No ratings yet
PT10 20 - Mobile - Pentesting - Preview
14 pages
Software Engineering - Chapter 7 - Detail Design - 1004486
No ratings yet
Software Engineering - Chapter 7 - Detail Design - 1004486
37 pages
Mathematical Logic through Python Yannai A. Gonczarowski instant download
No ratings yet
Mathematical Logic through Python Yannai A. Gonczarowski instant download
82 pages
Example of Summary of Findings in A Research Paper
No ratings yet
Example of Summary of Findings in A Research Paper
7 pages
ACE Truck ANSI X12 997
No ratings yet
ACE Truck ANSI X12 997
12 pages
Basic Settings for SAP EWM in SAP S_4HANA 1809 _ SAP Blogs
No ratings yet
Basic Settings for SAP EWM in SAP S_4HANA 1809 _ SAP Blogs
20 pages
Mr. Fourcan Karim Mazumder Faculty Dept. of Computer Science and Engineering
No ratings yet
Mr. Fourcan Karim Mazumder Faculty Dept. of Computer Science and Engineering
12 pages

Topic 4 W4 - Text Processing

Uploaded by

Topic 4 W4 - Text Processing

Uploaded by

Search Engines

Information Retrieval in Practice

All slides ©Addison Wesley, 2008

• How many pages contain all of the query terms?

• Assuming that terms occur independently

Collection size (N) is 25,205,179

Words with the Arabic root ktb

You might also like