Search Methods:
Keyword Search: Traditional search using exact word matching (TF-IDF, BM25)
Semantic Search: Understanding meaning using embedding models
Embedding Models:
Convert text to numerical vectors representing meaning
Local models (Sentence Transformers) vs. API models (OpenAI)
Demonstrated using the all-miniLM-L6-v2 model
Vector Databases:
Efficiently store and search embeddings
Introduced ChromaDB for learning and Pinecone for production
Document Chunking:
Breaking large documents into smaller, searchable pieces
Strategies: fixed-size chunks, sentence-based, paragraph-based
Importance of overlap to preserve context
Comments
0 B
|👍
/👎
0 B
|👍
/👎
0 B
|👍
/👎