joshualoehr/ngram-language-modelPublic

NotificationsYou must be signed in to change notification settings
Fork27
Star88

Python implementation of an N-gram language model with Laplace smoothing and sentence generation.

88 stars 27 forks Branches Tags Activity

You must be signed in to change notification settings

Folders and files

Name		Name	Last commit message	Last commit date
Latest commit History 5 Commits
data		data
README.md		README.md
language_model.py		language_model.py
preprocess.py		preprocess.py

Repository files navigation

N-Gram Language Model

Python implementation of an N-gram language model with Laplace smoothing and sentence generation.

Some NLTK functions are used (nltk.ngrams,nltk.FreqDist), but most everything is implemented by hand.

Note: theLanguageModel class expects to be given data which is already tokenized by sentences. If using the includedload_data function, thetrain.txt andtest.txt files should already be processed such that:

punctuation is removed
each sentence is on its own line

See thedata/ directory for examples.

Example output for a trigram model trained ondata/train.txt and tested againstdata/test.txt:

Loading 3-gram model...Vocabulary size: 23505Generating sentences......<s> <s> the company said it has agreed to sell its shares in a statement </s> (0.03163)<s> <s> he said the company also announced measures to boost its domestic economy and could be a long term debt </s> (0.01418)<s> <s> this is a major trade bill that would be the first quarter of 1987 </s> (0.02182)...Model perplexity: 51.555

The numbers in parentheses beside the generated sentences are the cumulative probabilities of those sentences occurring.

Usage info:

usage: N-gram Language Model [-h] --data DATA --n N [--laplace LAPLACE] [--num NUM]optional arguments:  -h, --help         show this help message and exit  --data DATA        Location of the data directory containing train.txt and test.txt  --n N              Order of N-gram model to create (i.e. 1 for unigram, 2 for bigram, etc.)  --laplace LAPLACE  Lambda parameter for Laplace smoothing (default is 0.01 -- use 1 for add-1 smoothing)  --num NUM          Number of sentences to generate (default 10)

Originally authored by Josh Loehr and Robin Cosbey, with slight modifications. Last edited Feb. 8, 2018.

About

Python implementation of an N-gram language model with Laplace smoothing and sentence generation.

Releases

No releases published

Packages

No packages published

Languages

Python100.0%

Movatterモバイル変換

Navigation Menu

Search code, repositories, users, issues, pull requests...

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Folders and files

Latest commit

History

Repository files navigation

N-Gram Language Model

About

Topics

Resources

Uh oh!

Stars

Watchers

Forks

Releases

Packages

Languages

Movatterモバイル変換

joshualoehr/ngram-language-model

Folders and files

Latest commit

History

Repository files navigation

N-Gram Language Model

About

Topics

Resources

Uh oh!

Stars

Watchers

Forks

Releases

Packages0

Languages

Packages