Get in Touch

Course Outline

Comprehensive training structure

  1. Introduction to NLP
    • Grasping NLP concepts
    • NLP frameworks
    • Commercial uses of NLP
    • Extracting data from the web
    • Utilizing diverse APIs to fetch text data
    • Managing text corpora, including content storage and associated metadata
    • Benefits of Python and an NLTK intensive session
  2. Practical Insights into Corpora and Datasets
    • The necessity of a corpus
    • Corpus examination
    • Categories of data attributes
    • Various file formats for corpora
    • Preparing datasets for NLP tasks
  3. Deciphering Sentence Structure
    • Essential NLP components
    • Natural language comprehension
    • Morphological analysis - stemming, words, tokens, and speech tags
    • Syntactic analysis
    • Semantic analysis
    • Managing ambiguity
  4. Text Data Preprocessing
    • Corpus - raw text
      • Sentence tokenization
      • Stemming raw text
      • Lemmatizing raw text
      • Eliminating stop words
    • Corpus - raw sentences
      • Word tokenization
      • Word lemmatization
    • Utilizing Term-Document/Document-Term matrices
    • Tokenizing text into n-grams and sentences
    • Customized and practical preprocessing
  5. Text Data Analysis
    • Fundamental NLP features
      • Parsers and parsing
      • POS tagging and taggers
      • Named entity recognition
      • N-grams
      • Bag of words
    • Statistical NLP features
      • Linear algebra concepts for NLP
      • Probabilistic theory for NLP
      • TF-IDF
      • Vectorization
      • Encoders and Decoders
      • Normalization
      • Probabilistic Models
    • Advanced feature engineering in NLP
      • Word2vec fundamentals
      • Word2vec model components
      • Logic behind the word2vec model
      • Extending the word2vec concept
      • Applications of the word2vec model
    • Case study: Applying bag of words for automatic text summarization using simplified and true Luhn's algorithms
  6. Document Clustering, Classification, and Topic Modeling
    • Document clustering and pattern discovery (hierarchical clustering, k-means, etc.)
    • Document comparison and classification using TFIDF, Jaccard, and cosine distance metrics
    • Document classification using Naïve Bayes and Maximum Entropy
  7. Identifying Key Text Elements
    • Dimensionality reduction: Principal Component Analysis, Singular Value Decomposition, and non-negative matrix factorization
    • Topic modeling and information retrieval via Latent Semantic Analysis
  8. Entity Extraction, Sentiment Analysis, and Advanced Topic Modeling
    • Positive vs. negative: Sentiment intensity
    • Item Response Theory
    • Part of speech tagging applications: Identifying people, places, and organizations in text
    • Advanced topic modeling: Latent Dirichlet Allocation
  9. Case Studies
    • Analyzing unstructured user reviews
    • Sentiment classification and visualization of Product Review Data
    • Mining search logs for usage patterns
    • Text classification
    • Topic modelling

Requirements

A solid grasp of NLP fundamentals and an understanding of how AI is applied in business contexts

 21 Hours

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories