Gemini Enterprise Agent Platform 生成式 AI 的配額和系統限制

本頁面提供各區域和模型的配額清單,並說明如何在 Google Cloud 控制台中查看及編輯配額。

調整後模型配額

微調模型推論與基礎模型共用配額。 微調模型推論沒有獨立配額。

嵌入限制

gemini-embedding-001 的要求受區域配額限制,而 gemini-embedding-2 的要求則受全域配額限制。
基礎模型 配額 指標
base_model: gemini-embedding 100,000,000 aiplatform.googleapis.com/embed_content_input_tokens_per_minute_per_base_model
base_model: gemini-embedding-2 200,000,000 aiplatform.googleapis.com/global_embed_content_input_tokens_per_minute_per_base_model
base_model: gemini-embedding-2 60,000 aiplatform.googleapis.com/global_embed_content_requests_per_minute_per_base_model

使用 predict API 要求 gemini-embedding-001 時,也須遵守下列配額:

基礎模型 配額 指標
base_model: gemini-embedding 100,000 aiplatform.googleapis.com/online_prediction_requests_per_base_model
base_model:不適用 30,000 aiplatform.googleapis.com/online_prediction_requests

Agent Runtime 配額

下列配額適用於每個區域中特定專案的 Agent Runtime
說明 配額 指標
每分鐘建立、刪除或更新 Agent Runtime 資源 10 aiplatform.googleapis.com/reasoning_engine_service_write_requests
每分鐘建立、刪除或更新 Agent Runtime 工作階段 100 aiplatform.googleapis.com/session_write_requests
每分鐘取得、列出或擷取 Agent Runtime 工作階段 10000 aiplatform.googleapis.com/session_read_requests
QueryStreamQuery Agent Runtime (每分鐘) 90 aiplatform.googleapis.com/reasoning_engine_service_query_requests
每分鐘將事件附加至 Agent Runtime 工作階段 300 aiplatform.googleapis.com/session_event_append_requests
Agent Runtime 資源數量上限 100 aiplatform.googleapis.com/reasoning_engine_service_entities
每分鐘建立、刪除或更新 Agent Runtime 記憶體資源 100 aiplatform.googleapis.com/memory_bank_write_requests
每分鐘從 Agent Runtime Memory Bank 取得、列出或擷取資料 300 aiplatform.googleapis.com/memory_bank_read_requests
沙箱環境 (程式碼執行) 每分鐘的執行要求數 1000 aiplatform.googleapis.com/sandbox_environment_execute_requests
每個區域的沙箱環境 (程式碼執行) 實體 1000 aiplatform.googleapis.com/sandbox_environment_entities
沙箱環境 (程式碼執行) 每分鐘寫入要求數 500 aiplatform.googleapis.com/sandbox_environment_write_requests
每分鐘的 A2A 代理程式 POST 要求,例如 sendMessagecancelTask 60 aiplatform.googleapis.com/a2a_agent_post_requests
每分鐘的 A2A 代理要求,例如 getTaskgetCard 600 aiplatform.googleapis.com/a2a_agent_get_requests
每分鐘使用 BidiStreamQuery API 的並行雙向即時連線數 10 aiplatform.googleapis.com/reasoning_engine_service_concurrent_query_requests

多模態輸入配額

特定專案在各個區域的 generateContentstreamGenerateContent 要求,適用下列多模態輸入內容配額。系統會根據基礎模型和解析度,分別強制執行各項配額。全域端點服務的要求也適用相同限制,並使用對應的..._global指標。
說明 配額 指標
每分鐘每個基礎模型和解析度的圖片輸入內容生成要求數 34,000,000 aiplatform.googleapis.com/generate_content_image_input_per_base_model_id_and_resolution
每分鐘每個基礎模型和解析度的影片輸入內容生成要求數 192,000,000 aiplatform.googleapis.com/generate_content_video_input_per_base_model_id_and_resolution
每分鐘以音訊輸入內容生成內容要求數 (依據基礎模型和解析度) 11,000,000 aiplatform.googleapis.com/generate_content_audio_input_per_base_model_id_and_resolution
每分鐘每個基礎模型和解析度的文件輸入內容生成要求數 1,200,000 aiplatform.googleapis.com/generate_content_document_input_per_base_model_id_and_resolution

如要提高這些配額的上限,請與 Google Cloud 帳戶團隊聯絡,因為無法透過 Google Cloud 控制台提高上限。

批次推論

所有區域的批次推論工作配額和限制都相同。

Gemini 模型並行批次推論工作限制

Gemini 模型沒有預先定義的批次推論配額限制。批次服務會提供大型共用資源集區的存取權,並根據模型即時可用性,以及所有客戶對該模型的需求,動態分配資源。如果模型容量已達上限,且有大量顧客處於活躍狀態,系統可能會將批次要求排入容量佇列。

非 Gemini 模型並行批次推論工作配額

下表列出並行批次推論工作數量的配額,不適用於 Gemini 模型:
配額
aiplatform.googleapis.com/textembedding_gecko_concurrent_batch_prediction_jobs 4
如果提交的工作數量超過分配的配額,系統會將工作放入佇列,並在配額容量可用時處理工作。

在 Google Cloud 控制台中查看及編輯配額

如要在 Google Cloud 控制台中查看及編輯配額,請按照下列步驟操作:
  1. 前往「配額與系統限制」頁面。
  2. 前往「配額與系統限制」

  3. 如要調整配額,請複製並貼上 Filter 中的屬性 aiplatform.googleapis.com/textembedding_gecko_concurrent_batch_prediction_jobs。按下「Enter」鍵。
  4. 按一下資料列末端的三點圖示,然後選取「編輯配額」
  5. 在窗格中輸入新的配額值,然後按一下「提交要求」

語意管理政策配額

下列配額適用於每個區域中特定專案的語意控管政策
說明 配額 指標
每分鐘取得或列出語意管理政策資源 600 aiplatform.googleapis.com/semantic_governance/policy_read_requests
每分鐘建立、更新或刪除語意管理政策資源 60 aiplatform.googleapis.com/semantic_governance/policy_write_requests
每分鐘取得語意管理政策引擎資源 600 aiplatform.googleapis.com/semantic_governance/engine_read_requests
每分鐘更新或取消佈建語意管理政策引擎資源 60 aiplatform.googleapis.com/semantic_governance/engine_write_requests
每個專案在每個位置的語意管理政策資源數。 1,000 aiplatform.googleapis.com/semantic_governance/policy_count

在 Google Cloud 控制台中查看及編輯配額

如要在 Google Cloud 控制台中查看及編輯配額,請按照下列步驟操作:
  1. 前往「配額與系統限制」頁面。
  2. 前往「配額與系統限制」

  3. 如要調整配額,請複製並貼上 Filter 中的屬性 aiplatform.googleapis.com/semantic_governance/policy_write_requests。按下「Enter」鍵。
  4. 按一下資料列末端的三點圖示,然後選取「編輯配額」
  5. 在窗格中輸入新的配額值,然後按一下「提交要求」

Gemini Enterprise Agent Platform 的 RAG 引擎

如要讓各項服務使用 RAG 引擎執行檢索增強生成 (RAG),必須遵守下列配額,配額的計算單位為每分鐘要求數 (RPM)。
服務 配額 指標
RAG Engine 資料管理 API 60 RPM VertexRagDataService requests per minute per region
RetrievalContexts 個 API 600 RPM VertexRagService retrieve requests per minute per region
base_model: textembedding-gecko 1,500 RPM Online prediction requests per base model per minute per region per base_model

您可以指定的額外篩選條件為 base_model: textembedding-gecko
以下限制適用於這類要求:
服務 限制 指標
並行 ImportRagFiles 要求 3 RPM VertexRagService concurrent import requests per region
每個 ImportRagFiles 要求的檔案數量上限 10,000 VertexRagService import rag files requests per region

Gen AI Evaluation Service

Gen AI Evaluation Service 會使用 Gemini 2.5 Flash 做為以模型為基礎的指標的預設評估模型。 以模型為基礎的指標單一評估要求,可能會導致對 Gen AI Evaluation Service 的多個基礎要求。系統會以機構層級計算各模型的用量,也就是說,凡是導向評估模型的要求 (用於模型推論和模型評估),都會計入模型的用量。下表列出 Gen AI Evaluation Service 和基礎評估模型配額:
要求配額 預設配額
每分鐘的 Gen AI Evaluation Service 要求數 每項專案每個區域 1,000 個要求
Gemini 處理量 視模型和計費方案而定
並行評估執行作業 每個區域每項專案的並行評估執行次數為 20 次

如果您在使用 Gen AI Evaluation Service 時收到配額相關錯誤,可能需要提出配額提高要求。詳情請參閱「查看及管理配額」。

限制
Gen AI Evaluation Service 服務要求逾時 60 秒

在新的專案中首次使用 Gen AI Evaluation Service 時,初始設定可能會延遲最多兩分鐘。如果第一次要求失敗,請稍候幾分鐘再試一次。後續的評估要求通常會在 60 秒內完成。

模型輸入和輸出權杖的上限取決於用來做為評估模型的模型。如需型號清單,請參閱 Google 型號

Gemini Enterprise Agent Platform Pipelines 配額

每個微調作業都會使用 Gemini Enterprise Agent Platform Pipelines。詳情請參閱「Agent Platform Pipelines 配額與限制」。

後續步驟

總覽

瞭解 Standard PayGo 方案。這個 Agent Platform 消費選項可讓您只為耗用的資源付費,無須預先承諾財務支出。

資源

與 Agent Platform 相關的配額和系統限制,不包括產品專屬的配額和系統限制。

總覽

瞭解 Google Cloud 如何限制 Google Cloud 雲端專案可使用的資源數量,以及配額如何適用於各種資源類型,包括硬體、軟體和網路元件。