8 种 RAG 架构,AI 工程师必懂(附适用场景)
1)Naive RAG
基于 Query 与文档向量的相似度进行检索。
适合事实查询、知识问答等简单场景,语义匹配即可满足需求。
2)Multimodal RAG(多模态 RAG)
同时检索文本、图片、音频等多种数据。
适用于跨模态任务,例如根据文字问题结合图片和文本进行回答。
3)HyDE(Hypothetical Document Embeddings)
当 Query 与目标文档语义差距较大时效果更好。
先让模型根据 Query 生成一篇“假想答案”,再对其进行向量化检索。
利用假想文档的语义表示找到更相关的真实文档。
4)Corrective RAG(校正型 RAG)
对检索结果进行二次验证,例如结合 Web Search 等可信来源。
在送入 LLM 前过滤、修正过时或错误信息,提高回答准确性。
5)Graph RAG
将检索结果组织成知识图谱,显式表示实体及其关系。
为 LLM 提供结构化上下文,增强推理能力,尤其适合复杂知识关联场景。
6)Hybrid RAG(混合 RAG)
将向量检索与图检索结合到同一流程中。
同时利用非结构化文本和结构化关系数据,获得更完整、更准确的答案。
7)Adaptive RAG(自适应 RAG)
根据问题复杂度动态选择检索策略。
简单问题直接检索,复杂问题自动拆分为多个子问题进行多轮检索与推理。
8)Agentic RAG(Agent RAG)
引入 AI Agent,通过规划、推理(如 ReAct、CoT)和长期记忆协调检索流程。
可以调用工具、API、多数据源,甚至组合多种 RAG 方法完成复杂任务。
大多数 RAG 架构的核心区别,都发生在检索阶段(Retrieval Time):决定检索什么、检索多少、如何组合信息。
但它们都建立在同一个前提上——索引(Indexing)已经做好了。
如果索引阶段切分的质量很差,无论采用哪种 RAG 架构,最终效果都会受到限制。优化索引质量,本身就是一个独立且非常重要的问题。
Akshay写了一种新的 Indexing 方法,声称可以做到:
将语料规模减少约 40 倍。
单次查询 Token 消耗降低约 3 倍。
向量检索相关性提升约 2.3 倍。
而且,它无需修改 Retrieval 算法、Reranker 或 Embedding 模型,只优化了 Indexing 本身。
8 RAG architectures for AI Engineers:
(explained with usage)
1) Naive RAG
- Retrieves documents purely based on vector similarity between the query embedding and stored embeddings.
- Works best for simple, fact-based queries where direct semantic matching suffices.
2) Multimodal RAG
- Handles multiple data types (text, images, audio, etc.) by embedding and retrieving across modalities.
- Ideal for cross-modal retrieval tasks like answering a text query with both text and image context.
3) HyDE (Hypothetical Document Embeddings)
- Queries are not semantically similar to documents.
- This technique generates a hypothetical answer document from the query before retrieval.
- Uses this generated document’s embedding to find more relevant real documents.
4) Corrective RAG
- Validates retrieved results by comparing them against trusted sources (e.g., web search).
- Ensures up-to-date and accurate information, filtering or correcting retrieved content before passing to the LLM.
5) Graph RAG
- Converts retrieved content into a knowledge graph to capture relationships and entities.
- Enhances reasoning by providing structured context alongside raw text to the LLM.
6) Hybrid RAG
- Combines dense vector retrieval with graph-based retrieval in a single pipeline.
- Useful when the task requires both unstructured text and structured relational data for richer answers.
7) Adaptive RAG
- Dynamically decides if a query requires a simple direct retrieval or a multi-step reasoning chain.
- Breaks complex queries into smaller sub-queries for better coverage and accuracy.
8) Agentic RAG
- Uses AI agents with planning, reasoning (ReAct, CoT), and memory to orchestrate retrieval from multiple sources.
- Best suited for complex workflows that require tool use, external APIs, or combining multiple RAG techniques.
Most architectures here involve some form of retrieval-time decision. But they all run on top of whatever was already indexed.
If that indexing step outputs messy chunks, every architecture inherits them. Improving it is a separate problem from the 8 above.
I wrote about a better unit for the indexing step. The technique:
- cuts corpus size by 40x.
- reduces tokens per query by 3x.
- improves vector search relevance by 2.3x.
And it doesn't alter the retrieval algorithm, the reranker, or the embedding model.
Read it in the article quoted below.
顯示更多