I'm building a Python algorithm that generates a random new page based on the words in an entire book and how frequently those words appear. I need to extract the PDF's text, identify each unique word, and count how many times it occurs. What tools or approach would work best for this? The generated page does not need to make grammatical sense.
3 Answers
A basic Python approach is to create a Counter and feed it the words as you read them. For example: `from collections import Counter; counts = Counter(text.lower().split())`. Then `counts['example']` gives the frequency of a word, and `counts.most_common()` lists words from most frequent to least frequent. For better results, handle punctuation and decide whether words such as “the” and “The” should be treated as the same.
You could also load the extracted text into a database and group by the normalized word with a count, but that is probably unnecessary for a single book. Python’s Counter is simpler and keeps both the unique-word list and frequencies in one object.
First convert the PDF into plain text with a tool such as pdftotext or a Python PDF library. Once you have the text, normalize it—usually by converting it to lowercase and removing or standardizing punctuation—then split it into words and count them with Python’s collections.Counter.

The important part is getting reliable text extraction first. Scanned PDFs may require OCR before a word counter can work.