# Introduction to Natural Language Processing (NLP)

## What is NLP and who uses it?

Natural Language Processing (NLP) encompasses the computational analysis and generation of human-readable language (as opposed to machine language).
Hodeghatta and Nayak (2023) show that NLP is not only of interest to linguists or literary scholars but also applied outside academia, e.g. in the field of Business Intelligence.
In fact, NLP technologies are embedded in many text-oriented (web) applications that consumers use on a regular basis, such as online shops, chatbots, or news and weather apps.

## Why is it difficult?

<blockquote>Natural language processing (NPL) is an extremely difficult task in computer science. Languages present a wide variety of problems that vary from language to language. Structuring or extracting meaningful information from free text represents a great solution, if done in the right manner. Previously, computer scientists broke a language into its grammatical forms, such as parts of speech, phrases, etc., using complex algorithms.
Today, deep learning is a key to performing the same exercises.</blockquote> (Goyal et. al., 2018, abstract)

## What role does text analysis play in NLP?

Text analysis as we perform it in "Machines of Knowledge" is just one element of NLP. NLP also includes text-to-speech and speech-to-text technologies, the generation of texts through language models, and translations between languages.
Hodeghatta and Nayak outline the broad range of NLP applications in their book.

## What tools are used for NLP in research?

NLP can be performed with ready-made tools like Voyant or the Dariah Topics Explorer, but their functionality is naturally limited and users can only manipulate the underlying algorithms to a small degree.
Working with programming languages such as Python or R offers more flexibility. Commonly used NLP packages that integrate with Python and R are [spacy], [NLTK] etc.
Janco et al. (2019) explain the use of spaCy specifically for digital humanities research.

### Cited works and further reading recommendations

- Goyal, P., Pandey, S., & Jain, K. (2018). Introduction to natural language processing and deep learning.
In P. Goyal, S. Pandey, & K. Jain (Eds.), Deep Learning for Natural Language Processing: Creating Neural Networks with Python (pp. 1–74). Apress. https://doi.org/10.1007/978-1-4842-3685-7_1
- Hodeghatta, U. R., & Nayak, U. (2023). Introduction to natural language processing. In U. R. Hodeghatta & U. Nayak (Eds.), Practical Business Analytics Using R and Python: Solve Business Problems Using a Data-driven Approach (pp. 541–599). Apress. https://doi.org/10.1007/978-1-4842-8754-5_15
- Janco, A., Bernstein, S., & Lassner, D. (2019). Introduction to natural language processing for dh research with spacy—A fast and accessible library that integrates modern machine learning technology (Version 2) [Dataset]. DataverseNL. https://doi.org/10.34894/K3M9UA
- Thompson, C. A. (2000). A brief introduction to natural language processing for non-linguists. In J. Cussens & S. Džeroski (Eds.), Learning Language in Logic (pp. 36–48). 
Springer. https://doi.org/10.1007/3-540-40030-3_2
- EARHART, A. E. (2015). Data and the Fragmented Text: Tools, Visualization, and Datamining or Is Bigger Better? In Traces of the Old, Uses of the New (pp. 90–116). University of Michigan Press; JSTOR. https://doi.org/10.2307/j.ctv65swvf.8
- Hardeniya, N. (2015). NLTK essentials: Build cool NLP and machine learning applications using NLTK and other Python libraries. https://search.ebscohost.com/login.aspx?direct=true&scope=site&db=nlebk&db=nlabk&AN=1044817
- Ignatow, G., & Mihalcea, R. F. (2017). Text mining: A guidebook for the social sciences. SAGE. http://bvbr.bib-bvb.de:8991/F?func=service&doc_library=BVB01&local_base=BVB01&doc_number=029075357&sequence=000004&line_number=0002&func_code=DB_RECORDS&service_type=MEDIA
- Ignatow, G., & Mihalcea, R. F. (2018). An introduction to text mining: Research design, data collection, and analysis. Sage. http://scans.hebis.de/HEBCGI/show.pl?40248591_toc.pdf
- Kuhn, J. (2019). Computational text analysis within the Humanities: How to combine working practices from the contributing fields? Lang Resources & Evaluation Language Resources and Evaluation, 53(4), 565–602.
- Lynch, T. L. (2015). Where the Machine Stops: Software as Reader and the Rise of New Literatures. Research in the Teaching of English, 49(3), 297–304.
- M. Reese, R., & Bhatia, A. (2018). Natural Language Processing with Java: Techniques for Building Machine Learning and Neural Network Models for NLP, 2nd Edition. Packt Publishing Ltd. https://public.ebookcentral.proquest.com/choice/publicfullrecord.aspx?p=5485028
- Moretti, F. (2013). Distant Reading. Verso.
- Nguyen, D., Liakata, M., DeDeo, S., Eisenstein, J., Mimno, D., Tromble, R., & Winters, J. (2020). How We Do Things With Words: Analyzing Text as Social and Cultural Data. Frontiers in Artificial Intelligence, 3, 62. https://doi.org/10.3389/frai.2020.00062
- Piotrowski, M. (2012). Natural Language Processing for Historical Texts. Morgan & Claypool Publishers.
- Ramsay, S. (2011). Reading Machines: Toward an Algorithmic Criticism. University of Illinois Press. https://www.press.uillinois.edu/books/catalog/75tms2pw9780252036415.html
- Wilcock, G. (2009). Introduction to linguistic annotation and text analytics. Morgan & Claypool Publishers. http://site.ebrary.com/id/10515692



