0% found this document useful (0 votes)

3 views49 pages

Words and Corpora J+M

The document discusses various aspects of text processing, including word tokenization, normalization, and the complexities of handling different languages and their structures. It highlights the importance of understanding types and tokens in corpora, as well as the methodologies for segmenting and normalizing text for natural language processing tasks. Additionally, it covers algorithms like Byte Pair Encoding for subword tokenization and the challenges of morphological parsing and stemming.

Uploaded by

John Kevin 12A22

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

3 views49 pages

Words and Corpora J+M

Uploaded by

John Kevin 12A22

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

You are on page 1/ 49

Words and Corpora

Basic Text
Processing
How many words in a sentence?
"I do uh main- mainly business data processing"
◦ Fragments, filled pauses
"Seuss’s cat in the hat is different from other cats!"
◦ Lemma: same stem, part of speech, rough word sense
◦ cat and cats = same lemma
◦ Wordform: the full inflected surface form
◦ cat and cats = different wordforms
How many words in a sentence?
they lay back on the San Francisco grass and looked at the stars
and their

Type: an element of the vocabulary.

Token: an instance of that type in running text.
How many?
◦ 15 tokens (or 14)
◦ 13 types (or 12) (or 11?)
How many words in a corpus?
N = number of tokens
V = vocabulary = set of types, |V| is size of vocabulary
Heaps Law = Herdan's Law = where often .67 < β < .75
i.e., vocabulary size grows with > square root of the number of word tokens

Tokens = N Types = |V|

Switchboard phone conversations 2.4 million 20 thousand
Shakespeare 884,000 31 thousand
COCA 440 million 2 million
Google N-grams 1 trillion 13+ million
Corpora
Words don't appear out of nowhere!
A text is produced by
• a specific writer(s),
• at a specific time,
• in a specific variety,
• of a specific language,
• for a specific function.
Corpora vary along dimensions like
◦ Language: 7097 languages in the world
◦ Variety, like African American Language varieties.
◦ AAE Twitter posts might include forms like "iont" (I don't)
◦ Code switching, e.g., Spanish/English, Hindi/English:
S/E: Por primera vez veo a @username actually being hateful! It was beautiful:)
[For the first time I get to see @username actually being hateful! it was beautiful:) ]
H/E: dost tha or ra- hega ... dont wory ... but dherya rakhe
[“he was and will remain a friend ... don’t worry ... but have faith”]
◦ Genre: newswire, fiction, scientific articles, Wikipedia
◦ Author Demographics: writer's age, gender, ethnicity, SES
Corpus datasheets
Gebru et al (2020), Bender and Friedman (2018)

Motivation:
• Why was the corpus collected?
• By whom?
• Who funded it?
Situation: In what situation was the text written?
Collection process: If it is a subsample how was it sampled? Was
there consent? Pre-processing?
+Annotation process, language variety, demographics, etc.
Words and Corpora
Basic Text
Processing
Word tokenization
Basic Text
Processing
Text Normalization

Every NLP task requires text normalization:

1. Tokenizing (segmenting) words
2. Normalizing word formats
3. Segmenting sentences
Space-based tokenization
A very simple way to tokenize
◦ For languages that use space characters between words
◦ Arabic, Cyrillic, Greek, Latin, etc., based writing systems
◦ Segment off a token between instances of spaces
Unix tools for space-based tokenization
◦ The "tr" command
◦ Inspired by Ken Church's UNIX for Poets
◦ Given a text file, output the word tokens and their frequencies
Simple Tokenization in UNIX
(Inspired by Ken Church’s UNIX for Poets.)
Given a text file, output the word tokens and their frequencies
tr -sc ’A-Za-z’ ’\n’ < shakes.txt Change all non-alpha to
newlines
| sort Sort in alphabetical order
| uniq –c
Merge and count each type

1945 A
72 AARON
19 ABBESS
5 ABBOT 25 Aaron
6 Abate
... ...
1 Abates
5 Abbess
6 Abbey
3 Abbot
.... …
The first step: tokenizing
tr -sc ’A-Za-z’ ’\n’ < shakes.txt | head

THE
SONNETS
by
William
Shakespeare
From
fairest
creatures
We
...
The second step: sorting
tr -sc ’A-Za-z’ ’\n’ < shakes.txt | sort | head

A
A
A
A
A
A
A
A
A
...
More counting
Merging upper and lower case
tr ‘A-Z’ ‘a-z’ < shakes.txt | tr –sc ‘A-Za-z’ ‘\n’ | sort | uniq –c

Sorting the counts

tr ‘A-Z’ ‘a-z’ < shakes.txt | tr –sc ‘A-Za-z’ ‘\n’ | sort | uniq –c | sort –n –r

23243 the
22225 i
18618 and
16339 to
15687 of
12780 a
12163 you What happened here?
10839 my
10005 in
8954 d
Issues in Tokenization
Can't just blindly remove punctuation:
◦ m.p.h., Ph.D., AT&T, cap’n
◦ prices ($45.55)
◦ dates (01/02/06)
◦ URLs (http://www.northwestern.edu)
◦ hashtags (#nlproc)
◦ email addresses ([email protected])
Clitic: a word that doesn't stand on its own
◦ "are" in we're, French "je" in j'ai, "le" in l'honneur
When should multiword expressions (MWE) be words?
◦ New York, rock ’n’ roll
Tokenization in NLTK
Bird, Loper and Klein (2009), Natural Language Processing with Python. O’Reilly
Tokenization in languages without
spaces
Many languages (like Chinese, Japanese, Thai) don't
use spaces to separate words!

How do we decide where the token boundaries

should be?
Word tokenization in Chinese
Chinese words are composed of characters called
"hanzi" (or sometimes just "zi")
Each one represents a meaning unit called a morpheme.
Each word has on average 2.4 of them.
But deciding what counts as a word is complex and not
agreed upon.
How to do word tokenization in
Chinese?
姚明进入总决赛 “Yao Ming reaches the finals”
3 words?
姚明进入总决赛
YaoMing reaches finals
5 words?
姚明进入总决赛
Yao Ming reaches overall finals
7 characters? (don't use words at all):
姚明进入总决赛
Yao Ming enter enter overall decision game
How to do word tokenization in
Chinese?
姚明进入总决赛 “Yao Ming reaches the finals”
3 words?
姚明进入总决赛
YaoMing reaches finals
5 words?
姚明进入总决赛
Yao Ming reaches overall finals
7 characters? (don't use words at all):
姚明进入总决赛
Yao Ming enter enter overall decision game
How to do word tokenization in
Chinese?
姚明进入总决赛 “Yao Ming reaches the finals”
3 words?
姚明进入总决赛
YaoMing reaches finals
5 words?
姚明进入总决赛
Yao Ming reaches overall finals
7 characters? (don't use words at all):
姚明进入总决赛
Yao Ming enter enter overall decision game
How to do word tokenization in
Chinese?
姚明进入总决赛 “Yao Ming reaches the finals”
3 words?
姚明进入总决赛
YaoMing reaches finals
5 words?
姚明进入总决赛
Yao Ming reaches overall finals
7 characters? (don't use words at all):
姚明进入总决赛
Yao Ming enter enter overall decision game
Word tokenization / segmentation
So in Chinese it's common to just treat each character
(zi) as a token.
• So the segmentation step is very simple
In other languages (like Thai and Japanese), more
complex word segmentation is required.
• The standard algorithms are neural sequence models
trained by supervised machine learning.
Word tokenization
Basic Text
Processing
Byte Pair Encoding
Basic Text
Processing
Another option for text tokenization
Instead of
• white-space segmentation
• single-character segmentation
Use the data to tell us how to tokenize.
Subword tokenization (because tokens can be parts
of words as well as whole words)
Subword tokenization
Three common algorithms:
◦ Byte-Pair Encoding (BPE) (Sennrich et al., 2016)
◦ Unigram language modeling tokenization (Kudo, 2018)
◦ WordPiece (Schuster and Nakajima, 2012)
All have 2 parts:
◦ A token learner that takes a raw training corpus and induces
a vocabulary (a set of tokens).
◦ A token segmenter that takes a raw test sentence and
tokenizes it according to that vocabulary
Byte Pair Encoding (BPE) token learner
Let vocabulary be the set of all individual characters
= {A, B, C, D,…, a, b, c, d….}
Repeat:
◦ Choose the two symbols that are most frequently
adjacent in the training corpus (say 'A', 'B')
◦ Add a new merged symbol 'AB' to the vocabulary
◦ Replace every adjacent 'A' 'B' in the corpus with 'AB'.
Until k merges have been done.
BPE token learner algorithm
Byte Pair Encoding (BPE) Addendum
Most subword algorithms are run inside
space-separated tokens.
So we commonly first add a special end-of-word
symbol '__' before space in training corpus
Next, separate into letters.
BPE token learner
Original (very fascinating🙄) corpus:
low low low low low lowest lowest newer newer newer
newer newer newer wider wider wider new new

Add end-of-word tokens, resulting in this vocabulary:

representation
BPE token learner

Merge e r to er
BPE

Merge er _ to er_
BPE

Merge n e to ne
BPE
The next merges are:
BPE token segmenter algorithm
On the test data, run each merge learned from the
training data:
◦ Greedily
◦ In the order we learned them
◦ (test frequencies don't play a role)
So: merge every e r to er, then merge er _ to er_, etc.
Result:
◦ Test set "n e w e r _" would be tokenized as a full word
◦ Test set "l o w e r _" would be two tokens: "low er_"
Properties of BPE tokens
Usually include frequent words
And frequent subwords
• Which are often morphemes like -est or –er
A morpheme is the smallest meaning-bearing unit of a
language
• unlikeliest has 3 morphemes un-, likely, and -est
Byte Pair Encoding
Basic Text
Processing
Word Normalization and
other issues
Basic Text
Processing
Word Normalization
Putting words/tokens in a standard format
◦ U.S.A. or USA
◦ uhhuh or uh-huh
◦ Fed or fed
◦ am, is, be, are
Case folding
Applications like IR: reduce all letters to lower case
◦ Since users tend to use lower case
◦ Possible exception: upper case in mid-sentence?
◦ e.g., General Motors
◦ Fed vs. fed
◦ SAIL vs. sail

For sentiment analysis, MT, Information extraction

◦ Case is helpful (US versus us is important)
Lemmatization

Represent all words as their lemma, their shared root

= dictionary headword form:
◦ am, are, is → be
◦ car, cars, car's, cars' → car
◦ Spanish quiero (‘I want’), quieres (‘you want’)
→ querer ‘want'
◦ He is reading detective stories
→ He be read detective story
Lemmatization is done by Morphological
Parsing
Morphemes:
◦ The small meaningful units that make up words
◦ Stems: The core meaning-bearing units
◦ Affixes: Parts that adhere to stems, often with grammatical
functions
Morphological Parsers:
◦ Parse cats into two morphemes cat and s
◦ Parse Spanish amaren (‘if in the future they would love’) into
morpheme amar ‘to love’, and the morphological features
3PL and future subjunctive.
Stemming
Reduce terms to stems, chopping off affixes crudely
This was not the map we
Thi wa not the map we
found in Billy Bones’s
found in Billi Bone s chest
chest, but an accurate
but an accur copi complet
copy, complete in all
in all thing name and
things-names and heights
height and sound with the
and soundings-with the
singl except of the red
single exception of the
cross and the written note
red crosses and the
.
written notes.
Porter Stemmer
Based on a series of rewrite rules run in series
◦ A cascade, in which output of each pass fed to next pass
Some sample rules:
Dealing with complex morphology is
necessary for many languages
◦ e.g., the Turkish word:
Uygarlastiramadiklarimizdanmissinizcasina
◦ `(behaving) as if you are among those whom we could not civilize’
◦ Uygar `civilized’ + las `become’
+ tir `cause’ + ama `not able’
+ dik `past’ + lar ‘plural’
+ imiz ‘p1pl’ + dan ‘abl’
+ mis ‘past’ + siniz ‘2pl’ + casina ‘as if’
Sentence Segmentation
!, ? mostly unambiguous but period “.” is very ambiguous
◦ Sentence boundary
◦ Abbreviations like Inc. or Dr.
◦ Numbers like .02% or 4.3
Common algorithm: Tokenize first: use rules or ML to
classify a period as either (a) part of the word or (b) a
sentence-boundary.
◦ An abbreviation dictionary can help
Sentence segmentation can then often be done by rules
based on this tokenization.
Word Normalization and
other issues
Basic Text
Processing

Spoken kannada class notes
No ratings yet
Spoken kannada class notes
44 pages
NLP Sem Answers (All)
No ratings yet
NLP Sem Answers (All)
124 pages
Corpora
No ratings yet
Corpora
48 pages
Lect 4 Words and Tokenizing
No ratings yet
Lect 4 Words and Tokenizing
24 pages
2 TextProc 2023
No ratings yet
2 TextProc 2023
35 pages
Regular Expression and BPE
No ratings yet
Regular Expression and BPE
68 pages
Week 2
No ratings yet
Week 2
90 pages
Text Proc
No ratings yet
Text Proc
55 pages
Text Preprocessing
No ratings yet
Text Preprocessing
59 pages
5 BASIC TEXT PROCESSING
No ratings yet
5 BASIC TEXT PROCESSING
6 pages
NLP_Week_02
No ratings yet
NLP_Week_02
55 pages
AI6122 Topic 1.2 - WordLevel
No ratings yet
AI6122 Topic 1.2 - WordLevel
63 pages
Lecture 2 NLP
No ratings yet
Lecture 2 NLP
27 pages
Introduction To NLP
No ratings yet
Introduction To NLP
68 pages
week_02_Tokenizers
No ratings yet
week_02_Tokenizers
36 pages
NLP_Week_02
No ratings yet
NLP_Week_02
54 pages
Apex Institute of Technology Natural Language Processing (20CST354)
No ratings yet
Apex Institute of Technology Natural Language Processing (20CST354)
43 pages
2 Text Processing
No ratings yet
2 Text Processing
58 pages
2 TextProc Mar 25 2021
No ratings yet
2 TextProc Mar 25 2021
71 pages
Week3
No ratings yet
Week3
15 pages
3.Word level analysis-tokenization stemming
No ratings yet
3.Word level analysis-tokenization stemming
8 pages
2.BasicTextProcessing NEW
No ratings yet
2.BasicTextProcessing NEW
39 pages
3.Chapter4_Lexical Representations
No ratings yet
3.Chapter4_Lexical Representations
36 pages
02 Textprocessingboth
No ratings yet
02 Textprocessingboth
46 pages
NATURAL LANGUAGE PROCESSING UNIT 1
No ratings yet
NATURAL LANGUAGE PROCESSING UNIT 1
16 pages
Kuhlmann - Introduction To Computational Linguistics (Slides) (2015)
100% (1)
Kuhlmann - Introduction To Computational Linguistics (Slides) (2015)
66 pages
PART B NOTES
No ratings yet
PART B NOTES
62 pages
Basic Text Processing: Regular Expressions
No ratings yet
Basic Text Processing: Regular Expressions
46 pages
Formalizing BPE Tokenization
No ratings yet
Formalizing BPE Tokenization
12 pages
Tokeniz prob!
No ratings yet
Tokeniz prob!
4 pages
Tokenization
No ratings yet
Tokenization
26 pages
Grading: Final Term: 40 % Term Paper: 30% Assignments and Quizzes: 30%
No ratings yet
Grading: Final Term: 40 % Term Paper: 30% Assignments and Quizzes: 30%
46 pages
Text preprocessing
No ratings yet
Text preprocessing
39 pages
Module 1 Nlp
No ratings yet
Module 1 Nlp
26 pages
NLP Digital Notes
No ratings yet
NLP Digital Notes
128 pages
NLP m2
No ratings yet
NLP m2
71 pages
NLP Lecture2 Text Pre Processing
No ratings yet
NLP Lecture2 Text Pre Processing
54 pages
2 Textprocessingboth
No ratings yet
2 Textprocessingboth
46 pages
Lecture 1 Text Preprocessing PDF
No ratings yet
Lecture 1 Text Preprocessing PDF
29 pages
Session1 2024_2025_ Natural Language Processing
No ratings yet
Session1 2024_2025_ Natural Language Processing
40 pages
Natural Language Processing (NLP) & Computational Linguistics
No ratings yet
Natural Language Processing (NLP) & Computational Linguistics
60 pages
CL_lec 6
No ratings yet
CL_lec 6
28 pages
2 TextProc 2023
No ratings yet
2 TextProc 2023
74 pages
Basic Text Processing: Regular Expressions
No ratings yet
Basic Text Processing: Regular Expressions
41 pages
NLP_Lecture_6_Week_3
No ratings yet
NLP_Lecture_6_Week_3
9 pages
lec2
No ratings yet
lec2
21 pages
AI_NLP
No ratings yet
AI_NLP
9 pages
All Practicals
No ratings yet
All Practicals
33 pages
C10_AI_UNIT 3_NLP_ HALF YEARLY
No ratings yet
C10_AI_UNIT 3_NLP_ HALF YEARLY
37 pages
03 Word Tokenization 14-26
No ratings yet
03 Word Tokenization 14-26
6 pages
Natural Language Processing 1
No ratings yet
Natural Language Processing 1
19 pages
Basic Text Process
No ratings yet
Basic Text Process
3 pages
3. text-processing
No ratings yet
3. text-processing
70 pages
Text Processing, Tokenization & Characteristics
100% (1)
Text Processing, Tokenization & Characteristics
89 pages
1009_nlp_ppt
No ratings yet
1009_nlp_ppt
31 pages
2-Regular expressions, Text Normalization, Edit Distance
No ratings yet
2-Regular expressions, Text Normalization, Edit Distance
42 pages
Unit 6 - AI (NLP)
No ratings yet
Unit 6 - AI (NLP)
37 pages
Natural Language Processing
No ratings yet
Natural Language Processing
6 pages
NLP_AI_X
No ratings yet
NLP_AI_X
6 pages
Chinese Writing: The 178 Most Common Characters from New HSK 1
From Everand
Chinese Writing: The 178 Most Common Characters from New HSK 1
Crystal Gong
5/5 (3)
Character Voices: A Workbook for Audiobook Narration: Narrated by the Author, #2
From Everand
Character Voices: A Workbook for Audiobook Narration: Narrated by the Author, #2
Renee Conoulty
5/5 (1)
Fig. 1. Relationship Between AI and Natural Language Processing Technology
No ratings yet
Fig. 1. Relationship Between AI and Natural Language Processing Technology
6 pages
Eyetracking Analysis of EAP Students' Regions of Interest in Computer-based Feedback on Grammar Usage & Organization
No ratings yet
Eyetracking Analysis of EAP Students' Regions of Interest in Computer-based Feedback on Grammar Usage & Organization
14 pages
01 - Big Pig On A Dig
No ratings yet
01 - Big Pig On A Dig
11 pages
How To Write A Commentary English Language Coursework
100% (1)
How To Write A Commentary English Language Coursework
4 pages
FRENCH THROUGH COMMUNICATIVE APPROACH
No ratings yet
FRENCH THROUGH COMMUNICATIVE APPROACH
2 pages
Benefits of Studying A Second Language
No ratings yet
Benefits of Studying A Second Language
21 pages
1 Closure Properties of Context-Free Languages: 1.1 Union
No ratings yet
1 Closure Properties of Context-Free Languages: 1.1 Union
11 pages
Conditionals 0-1-2
No ratings yet
Conditionals 0-1-2
12 pages
16. THPT Chuyên Phan Bội Châu - Nghệ an - Lần 1 - File Word Có Lời Giải Chi Tiết
No ratings yet
16. THPT Chuyên Phan Bội Châu - Nghệ an - Lần 1 - File Word Có Lời Giải Chi Tiết
22 pages
The Mongolian Script
100% (1)
The Mongolian Script
17 pages
2-Lexical Analysis Part1
No ratings yet
2-Lexical Analysis Part1
39 pages
2.1.1. Form: 2. Sentence Types (By Functions) 2.1. Declarative Sentences
No ratings yet
2.1.1. Form: 2. Sentence Types (By Functions) 2.1. Declarative Sentences
4 pages
Shannon Exam 01 - 27 - 24
No ratings yet
Shannon Exam 01 - 27 - 24
2 pages
Lectia I Timpurile Modului Indicativ
No ratings yet
Lectia I Timpurile Modului Indicativ
32 pages
Wish and If Only Handout
No ratings yet
Wish and If Only Handout
2 pages
La Unit Plan
No ratings yet
La Unit Plan
7 pages
Duolingo Guide
No ratings yet
Duolingo Guide
18 pages
How Economics Shaped Human Nature by Seth Roberts
No ratings yet
How Economics Shaped Human Nature by Seth Roberts
23 pages
English Language Scheme of Work Primary Year 1: Kurikulum Standard Sekolah Rendah
No ratings yet
English Language Scheme of Work Primary Year 1: Kurikulum Standard Sekolah Rendah
10 pages
English paper 1 note 3
No ratings yet
English paper 1 note 3
14 pages
TCC
No ratings yet
TCC
12 pages
UCLA General Education Master Course List: Foundations of Knowledge
No ratings yet
UCLA General Education Master Course List: Foundations of Knowledge
10 pages
Present Perfect
No ratings yet
Present Perfect
3 pages
UG NEP Syllabus
No ratings yet
UG NEP Syllabus
105 pages
Linguistics - Unit 3
No ratings yet
Linguistics - Unit 3
5 pages
Effective Communication For Public Speaking
No ratings yet
Effective Communication For Public Speaking
2 pages
Hieroglyphs Tutorial
100% (2)
Hieroglyphs Tutorial
21 pages
Kulikov Leonid The Vedic Yapresents Passives and Intransitiv
No ratings yet
Kulikov Leonid The Vedic Yapresents Passives and Intransitiv
1,024 pages
21 Ways To Write Better Songs
100% (3)
21 Ways To Write Better Songs
17 pages

Words and Corpora J+M

Uploaded by

Words and Corpora J+M

Uploaded by

Words and Corpora

Type: an element of the vocabulary.

Tokens = N Types = |V|

Every NLP task requires text normalization:

Sorting the counts

How do we decide where the token boundaries

Add end-of-word tokens, resulting in this vocabulary:

For sentiment analysis, MT, Information extraction

Represent all words as their lemma, their shared root

You might also like