RAG за 5 минут с использованием Vector DB и DeepSeek#

В этом руководстве представлено, как создать конвейер извлечения и генерации данных (RAG), используя Platform V Vector DB (далее - Vector DB) в качестве решения для хранения векторов и DeepSeek для обогащения семантических запросов. Конвейеры RAG улучшают ответы моделей больших языков (LLM), предоставляя контекстно релевантные данные.

Обзор#

В данном руководстве рассмотрены следующие шаги:

  1. Использование образца текста и преобразуем его в векторы с помощью FastEmbed.

  2. Отправка векторов в коллекцию Vector DB.

  3. Соединение Vector DB и DeepSeek в минимальный конвейер RAG.

  4. Задание разных вопросов DeepSeek и проверка точности ответов.

  5. Дополнение подсказок DeepSeek контентом, извлеченным из Vector DB.

  6. Оценка точности ответов до и после.

Архитектура#

deepseek-rag-architecture

Предварительные условия#

Убедитесь в наличии:

Настройка Vector DB#

pip install "qdrant-client[fastembed]>=1.14.1"

Vector DB будет выступать в роли базы знаний, предоставляющей информацию о контексте для подсказок, которые будут отправляться модели языка.

Можно получить бесплатный экземпляр облачного сервиса Vector DB по адресу http://cloud.qdrant.io.

QDRANT_URL = "https://xyz-example.eu-central.aws.cloud.qdrant.io:6333"
QDRANT_API_KEY = "<your-api-key>"

Создание экземпляра клиента Vector DB#

from qdrant_client import QdrantClient, models

client = QdrantClient(url=QDRANT_URL, api_key=QDRANT_API_KEY)

Построение базы знаний#

Vector DB использует векторные вложения фактов для обогащения исходной подсказки некоторым контекстом. Таким образом, необходимо хранить векторные вложения и факты, использованные для их создания.

Будет использоваться модель bge-base-en-v1.5 через библиотеку FastEmbed - легковесную, быструю библиотеку на Python для генерации вложений.

Клиент Vector DB предоставляет удобную интеграцию с FastEmbed, которая делает создание базы знаний очень простым делом.

Сначала нужно создать коллекцию, чтобы Vector DB знал, какие векторы он будет обрабатывать, а затем просто передаются необработанные документы, обернутые в models.Document, для вычисления и загрузки вложений.

collection_name = "knowledge_base"
model_name = "BAAI/bge-small-en-v1.5"
client.create_collection(
    collection_name=collection_name,
    vectors_config=models.VectorParams(size=384, distance=models.Distance.COSINE)
)
documents = [
    "Qdrant is a vector database & vector similarity search engine. It deploys as an API service providing search for the nearest high-dimensional vectors. With Qdrant, embeddings or neural network encoders can be turned into full-fledged applications for matching, searching, recommending, and much more!",
    "Docker helps developers build, share, and run applications anywhere — without tedious environment configuration or management.",
    "PyTorch is a machine learning framework based on the Torch library, used for applications such as computer vision and natural language processing.",
    "MySQL is an open-source relational database management system (RDBMS). A relational database organizes data into one or more data tables in which data may be related to each other; these relations help structure the data. SQL is a language that programmers use to create, modify and extract data from the relational database, as well as control user access to the database.",
    "NGINX is a free, open-source, high-performance HTTP server and reverse proxy, as well as an IMAP/POP3 proxy server. NGINX is known for its high performance, stability, rich feature set, simple configuration, and low resource consumption.",
    "FastAPI is a modern, fast (high-performance), web framework for building APIs with Python 3.7+ based on standard Python type hints.",
    "SentenceTransformers is a Python framework for state-of-the-art sentence, text and image embeddings. You can use this framework to compute sentence / text embeddings for more than 100 languages. These embeddings can then be compared e.g. with cosine-similarity to find sentences with a similar meaning. This can be useful for semantic textual similar, semantic search, or paraphrase mining.",
    "The cron command-line utility is a job scheduler on Unix-like operating systems. Users who set up and maintain software environments use cron to schedule jobs (commands or shell scripts), also known as cron jobs, to run periodically at fixed times, dates, or intervals.",
]
client.upsert(
    collection_name=collection_name,
    points=[
        models.PointStruct(
            id=idx,
            vector=models.Document(text=document, model=model_name),
            payload={"document": document},
        )
        for idx, document in enumerate(documents)
    ],
)

Настройка DeepSeek#

RAG изменяет способ взаимодействия с большими языками моделями. Задача ориентированная на знания, при которой модель может дать вымышленный ответ, превращается в задачу ориентированную на язык. Последняя предполагает, что модель должна извлекать значимую информацию и генерировать ответ. Большие языковые модели, если они реализованы правильно, должны выполнять задачи, ориентированные на язык.

Задание начинается с оригинальной подсказки, отправленной пользователем. Та же самая подсказка затем векторизуется и используется в качестве поискового запроса для наиболее релевантных фактов. Эти факты объединяются с исходной подсказкой для построения более длинной подсказки, содержащей больше информации.

Начните с простого вопроса напрямую:

prompt = """
What tools should I need to use to build a web service using vector embeddings for search?
"""

Использование API Deepseek требует предоставления ключа API. Его можно получить на платформе DeepSeek.

Теперь можно вызвать API завершения:

import requests
import json

# Fill the environmental variable with your own Deepseek API key
# See: https://platform.deepseek.com/api_keys
API_KEY = "<YOUR_DEEPSEEK_KEY>"

HEADERS = {
    "Authorization": f"Bearer {API_KEY}",
    "Content-Type": "application/json",
}


def query_deepseek(prompt):
    data = {
        "model": "deepseek-chat",
        "messages": [{"role": "user", "content": prompt}],
        "stream": False,
    }

    response = requests.post(
        "https://api.deepseek.com/chat/completions", headers=HEADERS, data=json.dumps(data)
    )

    if response.ok:
        result = response.json()
        return result["choices"][0]["message"]["content"]
    else:
        raise Exception(f"Error {response.status_code}: {response.text}")

и также запрос:

query_deepseek(prompt)

Ответ:

"Building a web service that uses vector embeddings for search involves several components, including data processing, embedding generation, storage, search, and serving the service via an API. Below is a list of tools and technologies you can use for each step:\n\n---\n\n### 1. **Data Processing**\n   - **Python**: For general data preprocessing and scripting.\n   - **Pandas**: For handling tabular data.\n   - **NumPy**: For numerical operations.\n   - **NLTK/Spacy**: For text preprocessing (tokenization, stemming, etc.).\n   - **LLM models**: For generating embeddings if you're using pre-trained models.\n\n---\n\n### 2. **Embedding Generation**\n   - **Pre-trained Models**:\n     - Embeddings (e.g., `text-embedding-ada-002`).\n     - Hugging Face Transformers (e.g., `Sentence-BERT`, `all-MiniLM-L6-v2`).\n     - Google Universal Sentence Encoder.\n   - **Custom Models**:\n     - TensorFlow/PyTorch: For training custom embedding models.\n   - **Libraries**:\n     - `sentence-transformers`: For generating sentence embeddings.\n     - `transformers`: For using Hugging Face models.\n\n---\n\n### 3. **Vector Storage**\n   - **Vector Databases**:\n     - Pinecone: Managed vector database for similarity search.\n     - Weaviate: Open-source vector search engine.\n     - Milvus: Open-source vector database.\n     - FAISS (Facebook AI Similarity Search): Library for efficient similarity search.\n     - Qdrant: Open-source vector search engine.\n     - Redis with RedisAI: For storing and querying vectors.\n   - **Traditional Databases with Vector Support**:\n     - PostgreSQL with pgvector extension.\n     - Elasticsearch with dense vector support.\n\n---\n\n### 4. **Search and Retrieval**\n   - **Similarity Search Algorithms**:\n     - Cosine similarity, Euclidean distance, or dot product for comparing vectors.\n   - **Libraries**:\n     - FAISS: For fast nearest-neighbor search.\n     - Annoy (Approximate Nearest Neighbors Oh Yeah): For approximate nearest neighbor search.\n   - **Vector Databases**: Most vector databases (e.g., Pinecone, Weaviate) come with built-in search capabilities.\n\n---\n\n### 5. **Web Service Framework**\n   - **Backend Frameworks**:\n     - Flask/Django/FastAPI (Python): For building RESTful APIs.\n     - Node.js/Express: If you prefer JavaScript.\n   - **API Documentation**:\n     - Swagger/OpenAPI: For documenting your API.\n   - **Authentication**:\n     - OAuth2, JWT: For securing your API.\n\n---\n\n### 6. **Deployment**\n   - **Containerization**:\n     - Docker: For packaging your application.\n   - **Orchestration**:\n     - Kubernetes: For managing containers at scale.\n   - **Cloud Platforms**:\n     - AWS (EC2, Lambda, S3).\n     - Google Cloud (Compute Engine, Cloud Functions).\n     - Azure (App Service, Functions).\n   - **Serverless**:\n     - AWS Lambda, Google Cloud Functions, or Vercel for serverless deployment.\n\n---\n\n### 7. **Monitoring and Logging**\n   - **Monitoring**:\n     - Prometheus + Grafana: For monitoring performance.\n   - **Logging**:\n     - ELK Stack (Elasticsearch, Logstash, Kibana).\n     - Fluentd.\n   - **Error Tracking**:\n     - Sentry.\n\n---\n\n### 8. **Frontend (Optional)**\n   - **Frontend Frameworks**:\n     - React, Vue.js, or Angular: For building a user interface.\n   - **Libraries**:\n     - Axios: For making API calls from the frontend.\n\n---\n\n### Example Workflow\n1. Preprocess your data (e.g., clean text, tokenize).\n2. Generate embeddings using a pre-trained model (e.g., Hugging Face).\n3. Store embeddings in a vector database (e.g., Pinecone or FAISS).\n4. Build a REST API using FastAPI or Flask to handle search queries.\n5. Deploy the service using Docker and Kubernetes or a serverless platform.\n6. Monitor and scale the service as needed.\n\n---\n\n### Example Tools Stack\n- **Embedding Generation**: Hugging Face `sentence-transformers`.\n- **Vector Storage**: Pinecone or FAISS.\n- **Web Framework**: FastAPI.\n- **Deployment**: Docker + AWS/GCP.\n\nBy combining these tools, you can build a scalable and efficient web service for vector embedding-based search."

Расширение подсказки#

Хотя оригинальный ответ звучит правдоподобно, но он неправильно ответил на вопрос. Вместо этого он дал общее описание стека приложений. Чтобы улучшить результаты, обогащение исходной подсказки описаниями доступных инструментов кажется одним из возможных вариантов. Следует воспользоваться семантической базой знаний, чтобы дополнить подсказку описаниями различных технологий!

results = client.query_points(
    collection_name=collection_name,
    query=models.Document(text=prompt, model=model_name),
    limit=3,
)
results

Ответ:

QueryResponse(points=[
    ScoredPoint(id=0, version=0, score=0.67437416, payload={'document': 'Qdrant is a vector database & vector similarity search engine. It deploys as an API service providing search for the nearest high-dimensional vectors. With Qdrant, embeddings or neural network encoders can be turned into full-fledged applications for matching, searching, recommending, and much more!'}, vector=None, shard_key=None, order_value=None), 
    ScoredPoint(id=6, version=0, score=0.63144326, payload={'document': 'SentenceTransformers is a Python framework for state-of-the-art sentence, text and image embeddings. You can use this framework to compute sentence / text embeddings for more than 100 languages. These embeddings can then be compared e.g. with cosine-similarity to find sentences with a similar meaning. This can be useful for semantic textual similar, semantic search, or paraphrase mining.'}, vector=None, shard_key=None, order_value=None), 
    ScoredPoint(id=5, version=0, score=0.6064749, payload={'document': 'FastAPI is a modern, fast (high-performance), web framework for building APIs with Python 3.7+ based on standard Python type hints.'}, vector=None, shard_key=None, order_value=None)
])

Для выполнения семантического поиска по набору описаний инструментов была использована исходная подсказка. Теперь можно использовать эти описания для дополнения подсказки и создания большего контекста.

context = "\n".join(r.payload['document'] for r in results.points)
context

Ответ:

'Qdrant is a vector database & vector similarity search engine. It deploys as an API service providing search for the nearest high-dimensional vectors. With Qdrant, embeddings or neural network encoders can be turned into full-fledged applications for matching, searching, recommending, and much more!\nFastAPI is a modern, fast (high-performance), web framework for building APIs with Python 3.7+ based on standard Python type hints.\nPyTorch is a machine learning framework based on the Torch library, used for applications such as computer vision and natural language processing.'

Создайте метаподсказку, комбинацию предполагаемой роли LLM, исходного вопроса и результатов семантического поиска, которая заставит LLM использовать предоставленный контекст.

Выполняя это задача эффективно превращается из ориентированной на знания, в задачу, ориентированную на язык, что должно уменьшить вероятность появления галлюцинаций. Это также должно сделать ответ звучащим более уместным.

metaprompt = f"""
You are a software architect. 
Answer the following question using the provided context. 
If you can't find the answer, do not pretend you know it, but answer "I don't know".

Question: {prompt.strip()}

Context: 
{context.strip()}

Answer:
"""

# Look at the full metaprompt
print(metaprompt)

Ответ:

You are a software architect. 
Answer the following question using the provided context. 
If you can't find the answer, do not pretend you know it, but answer "I don't know".
    
Question: What tools should I need to use to build a web service using vector embeddings for search?
    
Context: 
Qdrant is a vector database & vector similarity search engine. It deploys as an API service providing search for the nearest high-dimensional vectors. With Qdrant, embeddings or neural network encoders can be turned into full-fledged applications for matching, searching, recommending, and much more!
FastAPI is a modern, fast (high-performance), web framework for building APIs with Python 3.7+ based on standard Python type hints.
PyTorch is a machine learning framework based on the Torch library, used for applications such as computer vision and natural language processing.
    
Answer:

Текущая подсказка гораздо длиннее, так как были применены несколько стратегий, чтобы еще больше улучшить ответы:

  1. LLM выполняет роль архитектора программного обеспечения.

  2. Предоставляется больше контекста для ответа на вопрос.

  3. Если контекст не содержит полезной информации, модель не должна придумывать ответ.

Проверка#

Вопрос:

query_deepseek(metaprompt)

Ответ:

'To build a web service using vector embeddings for search, you can use the following tools:\n\n1. **Qdrant**: As a vector database and similarity search engine, Qdrant will handle the storage and retrieval of high-dimensional vectors. It provides an API service for searching and matching vectors, making it ideal for applications that require vector-based search functionality.\n\n2. **FastAPI**: This web framework is perfect for building the API layer of your web service. It is fast, easy to use, and based on Python type hints, which makes it a great choice for developing the backend of your service. FastAPI will allow you to expose endpoints that interact with Qdrant for vector search operations.\n\n3. **PyTorch**: If you need to generate vector embeddings from your data (e.g., text, images), PyTorch can be used to create and train neural network models that produce these embeddings. PyTorch is a powerful machine learning framework that supports a wide range of applications, including natural language processing and computer vision.\n\n### Summary:\n- **Qdrant** for vector storage and search.\n- **FastAPI** for building the web service API.\n- **PyTorch** for generating vector embeddings (if needed).\n\nThese tools together provide a robust stack for building a web service that leverages vector embeddings for search functionality.'

Тестирование конвейера RAG#

Используя предоставленный семантический контекст, модель лучше справляется с ответом на поставленные вопросы. Заключите RAG в функцию, чтобы ее можно было легко вызывать для разных подсказок:

def rag(question: str, n_points: int = 3) -> str:
    results = client.query_points(
        collection_name=collection_name,
        query=models.Document(text=question, model=model_name),
        limit=n_points,
    )

    context = "\n".join(r.payload["document"] for r in results.points)

    metaprompt = f"""
    You are a software architect. 
    Answer the following question using the provided context. 
    If you can't find the answer, do not pretend you know it, but only answer "I don't know".

    Question: {question.strip()}

    Context: 
    {context.strip()}

    Answer:
    """

    return query_deepseek(metaprompt)

Теперь проще задавать широкий спектр вопросов.

Вопрос:

rag("What can the stack for a web api look like?")

Ответ:

'The stack for a web API can include the following components based on the provided context:\n\n1. **Web Framework**: FastAPI can be used as the web framework for building the API. It is modern, fast, and leverages Python type hints for better development and performance.\n\n2. **Reverse Proxy/Web Server**: NGINX can be used as a reverse proxy or web server to handle incoming HTTP requests, load balancing, and serving static content. It is known for its high performance and low resource consumption.\n\n3. **Containerization**: Docker can be used to containerize the application, making it easier to build, share, and run the API consistently across different environments without worrying about configuration issues.\n\nThis stack provides a robust, scalable, and efficient setup for building and deploying a web API.'

Вопрос:

rag("Where is the nearest grocery store?")

Ответ:

"I don't know. The provided context does not contain any information about the location of the nearest grocery store."

Модель теперь способна:

  1. Пользоваться знаниями, хранящимися в хранилище векторов.

  2. Отвечать на основе предоставленного контекста, что она не может предоставить ответ.

Заключение#

В данном руководстве был продемонстрирован полезный механизм снижения рисков возникновения иллюзий в моделях больших языков.