本頁面說明如何將 RAG 引擎連結至 Gemini Enterprise Agent Platform 向量搜尋。
您也可以使用 RAG Engine with Agent Platform 向量搜尋 筆記本逐步操作。
設定 Agent Platform 向量搜尋
Agent Platform Vector Search 採用 Google 研究開發的向量搜尋技術,透過向量搜尋,您可以運用為 Google 搜尋、YouTube 和 Google Play 等 Google 產品奠定基礎的相同基礎架構。
如要與 RAG Engine 整合,必須使用空白的向量搜尋索引。
設定 Vertex AI SDK
如要設定 Vertex AI SDK,請參閱「設定」。
建立向量搜尋索引
如要建立與 RAG Corpus 相容的向量搜尋索引,該索引必須符合下列條件:
IndexUpdateMethod必須為STREAM_UPDATE,請參閱建立串流索引。距離度量類型必須明確設為下列其中一項:
DOT_PRODUCT_DISTANCECOSINE_DISTANCE
向量的維度必須與您打算在 RAG 語料庫中使用的嵌入模型一致。您可以根據所選項目調整其他參數,這些項目會決定是否可以調整額外參數。
Python
如要瞭解如何安裝或更新 Vertex AI SDK for Python,請參閱「安裝 Vertex AI SDK for Python」。 詳情請參閱 Python API 參考文件。
建立向量搜尋索引端點
RAG Engine 支援公開端點。
Python
如要瞭解如何安裝或更新 Vertex AI SDK for Python,請參閱「安裝 Vertex AI SDK for Python」。 詳情請參閱 Python API 參考文件。
將索引部署至索引端點
執行最鄰近搜尋前,必須先將索引部署至索引端點。
Python
如要瞭解如何安裝或更新 Vertex AI SDK for Python,請參閱「安裝 Vertex AI SDK for Python」。 詳情請參閱 Python API 參考文件。
如果是首次將索引部署至索引端點,系統會自動建構並啟動後端,大約需要 30 分鐘,索引才能儲存。首次部署後,索引會在幾秒內準備就緒。如要查看索引部署狀態,請開啟 Vector Search Console,選取「索引端點」分頁,然後選擇索引端點。
找出索引和索引端點的資源名稱,格式如下:
projects/${PROJECT_NUMBER}/locations/${LOCATION_ID}/indexes/${INDEX_ID}projects/${PROJECT_NUMBER}/locations/${LOCATION_ID}/indexEndpoints/${INDEX_ENDPOINT_ID}。
在 RAG Engine 中使用 Agent Platform 向量搜尋
設定向量搜尋執行個體後,請按照本節的步驟,將向量搜尋執行個體設為 RAG 應用程式的向量資料庫。
設定向量資料庫,建立 RAG 語料庫
建立 RAG 語料庫時,請只指定完整的 INDEX_ENDPOINT_NAME 和 INDEX_NAME。請務必使用索引和索引端點資源名稱的數字 ID。系統會建立 RAG 語料庫,並自動與向量搜尋索引建立關聯。系統會根據條件執行驗證。如未符合任何一項規定,系統就會拒絕要求。
Python
在試用這個範例之前,請先按照「使用用戶端程式庫的 Agent Platform 快速入門導覽課程」中的 Python 設定說明操作。
如要向 Agent Platform 進行驗證,請設定應用程式預設憑證。 詳情請參閱「為本機開發環境設定驗證機制」。
REST
# TODO(developer): Update and un-comment the following lines:
# CORPUS_DISPLAY_NAME = "YOUR_CORPUS_DISPLAY_NAME"
# Full index/indexEndpoint resource name
# Index: projects/${PROJECT_NUMBER}/locations/${LOCATION_ID}/indexes/${INDEX_ID}
# IndexEndpoint: projects/${PROJECT_NUMBER}/locations/${LOCATION_ID}/indexEndpoints/${INDEX_ENDPOINT_ID}
# INDEX_RESOURCE_NAME = "YOUR_INDEX_ENDPOINT_RESOURCE_NAME"
# INDEX_NAME = "YOUR_INDEX_RESOURCE_NAME"
# Call CreateRagCorpus API to create a new RagCorpus
curl -X POST -H "Authorization: Bearer $(gcloud auth print-access-token)" -H "Content-Type: application/json" https://${LOCATION_ID}-aiplatform.googleapis.com/v1beta1/projects/${PROJECT_NUMBER}/locations/${LOCATION_ID}/ragCorpora -d '{
"display_name" : '\""${CORPUS_DISPLAY_NAME}"\"',
"rag_vector_db_config" : {
"vertex_vector_search": {
"index":'\""${INDEX_NAME}"\"'
"index_endpoint":'\""${INDEX_ENDPOINT_NAME}"\"'
}
}
}'
# Call ListRagCorpora API to verify the RagCorpus is created successfully
curl -sS -X GET \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
"https://${LOCATION_ID}-aiplatform.googleapis.com/v1beta1/projects/${PROJECT_NUMBER}/locations/${LOCATION_ID}/ragCorpora"
使用 RAG API 匯入檔案
使用 ragFiles.import API 方法,將檔案從 Cloud Storage 或 Google 雲端硬碟匯入向量搜尋索引。檔案會嵌入並儲存在向量搜尋索引中。
REST
# TODO(developer): Update and uncomment the following lines:
# RAG_CORPUS_ID = "your-rag-corpus-id"
#
# Google Cloud Storage bucket/file location.
# For example, "gs://rag-fos-test/"
# GCS_URIS= "your-gcs-uris"
# Call ImportRagFiles API to embed files and store in the BigQuery table
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
https://us-central1-aiplatform.googleapis.com/v1beta1/projects/${PROJECT_NUMBER}/locations/us-central1/ragCorpora/${RAG_CORPUS_ID}/ragFiles:import \
-d '{
"import_rag_files_config": {
"gcs_source": {
"uris": '\""${GCS_URIS}"\"'
},
"rag_file_chunking_config": {
"chunk_size": 512
}
}
}'
# Call ListRagFiles API to verify the files are imported successfully
curl -X GET \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
https://us-central1-aiplatform.googleapis.com/v1beta1/projects/${PROJECT_NUMBER}/locations/us-central1/ragCorpora/${RAG_CORPUS_ID}/ragFiles
Python
如要瞭解如何安裝或更新 Vertex AI SDK for Python,請參閱「安裝 Vertex AI SDK for Python」。 詳情請參閱 Python API 參考文件。
使用 RAG API 擷取相關背景資訊
檔案匯入完成後,您可以使用 RetrieveContexts API,從向量搜尋索引擷取相關內容。
REST
# TODO(developer): Update and uncomment the following lines:
# RETRIEVAL_QUERY="your-retrieval-query"
#
# Full RAG corpus resource name
# Format:
# "projects/${PROJECT_NUMBER}/locations/us-central1/ragCorpora/${RAG_CORPUS_ID}"
# RAG_CORPUS_RESOURCE="your-rag-corpus-resource"
# Call RetrieveContexts API to retrieve relevant contexts
curl -X POST \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
https://us-central1-aiplatform.googleapis.com/v1beta1/projects/${PROJECT_NUMBER}/locations/us-central1:retrieveContexts \
-d '{
"vertex_rag_store": {
"rag_resources": {
"rag_corpus": '\""${RAG_CORPUS_RESOURCE}"\"',
},
},
"query": {
"text": '\""${RETRIEVAL_QUERY}"\"',
"similarity_top_k": 10
}
}'
Python
如要瞭解如何安裝或更新 Vertex AI SDK for Python,請參閱「安裝 Vertex AI SDK for Python」。 詳情請參閱 Python API 參考文件。
使用 Agent Platform Gemini API 生成內容
如要使用 Gemini 模型生成內容,請呼叫 Agent Platform GenerateContent API。在要求中指定 RAG_CORPUS_RESOURCE,API 就會自動從向量搜尋索引擷取資料。
REST
# TODO(developer): Update and uncomment the following lines:
# MODEL_ID=gemini-2.5-flash
# GENERATE_CONTENT_PROMPT="your-generate-content-prompt"
# GenerateContent with contexts retrieved from the FeatureStoreOnline index
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" https://us-central1-aiplatform.googleapis.com/v1beta1/projects/${PROJECT_NUMBER}/locations/us-central1/publishers/google/models/${MODEL_ID}:generateContent \
-d '{
"contents": {
"role": "user",
"parts": {
"text": '\""${GENERATE_CONTENT_PROMPT}"\"'
}
},
"tools": {
"retrieval": {
"vertex_rag_store": {
"rag_resources": {
"rag_corpus": '\""${RAG_CORPUS_RESOURCE}"\"',
},
"similarity_top_k": 8,
"vector_distance_threshold": 0.32
}
}
}
}'
Python
如要瞭解如何安裝或更新 Vertex AI SDK for Python,請參閱「安裝 Vertex AI SDK for Python」。 詳情請參閱 Python API 參考文件。