Language Variation over time(NLP)

  • Tech Stack: numpy, NLTK, pytorch, jupyter notebook, scikit-learn
  • Google Drive URL: Project Link

Aiming to use language engineering methods to classify documents to the decade on which they were written. We use the TF-IDF method of representing documents, and train a multinomial Logistic Regression model to classify documents from the historic dataset “Svenska partiprogram och valmanifest” to the decade in which they were written.