The goal of this project is to develop and critically assess methods for detecting word blends among frequent terms in Twitter data, and to express the knowledge that you have gained about this task in a short research paper. Twitter users use language innovatively, and coining new terms by blending two existing words is a common phenomenon, known as lexical blending. Consider the following examples:
Component 1 Component 2 Blend word
Britain exit Brexit
spoon fork spork
breakfast lunch brunch
You will detect occurrences of blend words among a pre-processed list of tokens from a Twitter dataset, using a reference set of English words from a dictionary, and using methods for approximate string matching as encountered in the lectures. We will also provide you with a set of tweets the token list was extracted from, which you may (but are not expected to) use. You will evaluate the output of your algorithm(s) against a list of true word blends. The project aims to reinforce concepts in approximate matching and evaluation, and to strengthen your skills in data analysis and problem-solving. The goal of this assignment is not to develop a system that achieves near-perfect precision (in fact, this is impossible – we are developing knowledge technologies after all!).
Deliverables:
1. One or more programs, implemented in the programming language(s) of your choice, which must:
- Process the data input file(s), to identify word blend candidates
- Identify word blend candidates among a set of tokens, with the help of a reference collection of tokens (dictionary)
- Evaluate the matches, with respect to the list of true word blends, using one or more evaluation metrics
Assume there are five files:
1. blends.txt
2. candidates.txt
3. dict.txt
4. tweets.txt
5. wordforms.txt