feat: OSVauco GCP Agent Master Repo — Phase 5 initial files (docs, env, readme)
This commit is contained in:
parent
ed72a2691d
commit
d086e9ed80
43
.env.example
Normal file
43
.env.example
Normal file
|
|
@ -0,0 +1,43 @@
|
||||||
|
# .env.example — Copy to .env and fill in your values
|
||||||
|
# Usage: cp .env.example .env && source .env
|
||||||
|
# NEVER commit .env to version control
|
||||||
|
|
||||||
|
# ── Required ─────────────────────────────────────────────────
|
||||||
|
export PROJECT_ID="your-gcp-project-id"
|
||||||
|
export REGION="us-central1"
|
||||||
|
export BILLING_ACCOUNT_ID="XXXXXX-XXXXXX-XXXXXX"
|
||||||
|
|
||||||
|
# ── Service Accounts ─────────────────────────────────────────
|
||||||
|
export AGENT_SA="agent-runner@${PROJECT_ID}.iam.gserviceaccount.com"
|
||||||
|
|
||||||
|
# ── Cloud Run ────────────────────────────────────────────────
|
||||||
|
export CLOUD_RUN_SERVICE="gcp-orchestrator"
|
||||||
|
export ARTIFACT_REPO="agent-images"
|
||||||
|
|
||||||
|
# ── VPC / Security ───────────────────────────────────────────
|
||||||
|
export VPC_NETWORK="default"
|
||||||
|
export CMEK_KEY_RING="agent-keyring"
|
||||||
|
export CMEK_KEY_NAME="agent-key"
|
||||||
|
|
||||||
|
# ── Vertex AI RAG Engine ─────────────────────────────────────
|
||||||
|
export RAG_CORPUS_DISPLAY_NAME="gcp-agent-corpus"
|
||||||
|
export RAG_CORPUS_NAME=""
|
||||||
|
|
||||||
|
# ── Memory Bank ──────────────────────────────────────────────
|
||||||
|
export MEMORY_INSTANCE_DISPLAY_NAME="${PROJECT_ID}-memory-bank"
|
||||||
|
export MEMORY_BANK_INSTANCE=""
|
||||||
|
|
||||||
|
# ── CI/CD ────────────────────────────────────────────────────
|
||||||
|
export GITHUB_OWNER="your-github-username"
|
||||||
|
export GITHUB_REPO="GCP_Agent_Master_Repo"
|
||||||
|
|
||||||
|
# ── Observability ────────────────────────────────────────────
|
||||||
|
export ALERT_EMAIL="your-email@example.com"
|
||||||
|
|
||||||
|
# ── Agent Models ─────────────────────────────────────────────
|
||||||
|
export ORCHESTRATOR_MODEL="gemini-2.0-flash"
|
||||||
|
export SUBAGENT_MODEL="gemini-2.0-flash"
|
||||||
|
|
||||||
|
# ── Optional: LiteLLM ────────────────────────────────────────
|
||||||
|
# export OPENAI_API_KEY=""
|
||||||
|
# export ANTHROPIC_API_KEY=""
|
||||||
82
docs/Cost_Management_Rules.md
Normal file
82
docs/Cost_Management_Rules.md
Normal file
|
|
@ -0,0 +1,82 @@
|
||||||
|
# Cost Management Rules
|
||||||
|
**OSVauco-NMTMD-GCOS | GCP Agent Master Repo**
|
||||||
|
|
||||||
|
## 1. Budget Alerts
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gcloud billing budgets create \
|
||||||
|
--billing-account="${BILLING_ACCOUNT_ID}" \
|
||||||
|
--display-name="Agent Platform Budget" \
|
||||||
|
--budget-amount=500USD \
|
||||||
|
--threshold-rule=percent=0.5,basis=CURRENT_SPEND \
|
||||||
|
--threshold-rule=percent=0.8,basis=CURRENT_SPEND \
|
||||||
|
--threshold-rule=percent=1.0,basis=CURRENT_SPEND
|
||||||
|
```
|
||||||
|
|
||||||
|
## 2. Resource Labeling
|
||||||
|
|
||||||
|
All GCP resources MUST carry:
|
||||||
|
|
||||||
|
| Label Key | Example | Purpose |
|
||||||
|
|-----------|---------|--------|
|
||||||
|
| env | prod/dev/staging | Environment |
|
||||||
|
| team | osvaucoe | Cost attribution |
|
||||||
|
| cost-center | ai-platform | Finance |
|
||||||
|
| agent | orchestrator/rag | Per-agent tracking |
|
||||||
|
|
||||||
|
## 3. Cloud Run Cost Controls
|
||||||
|
|
||||||
|
```bash
|
||||||
|
gcloud run deploy ${SERVICE} \
|
||||||
|
--max-instances=10 \
|
||||||
|
--concurrency=80 \
|
||||||
|
--timeout=300
|
||||||
|
```
|
||||||
|
|
||||||
|
- CPU throttling on for non-latency-critical agents
|
||||||
|
- min-instances=1 only for latency-sensitive services
|
||||||
|
|
||||||
|
## 4. Vertex AI Cost Controls
|
||||||
|
|
||||||
|
- Use gemini-2.0-flash for high-volume sub-agent calls
|
||||||
|
- Context caching for repeated prompts (up to 75% savings)
|
||||||
|
- max_output_tokens=2048 for most tasks
|
||||||
|
- Monitor: aiplatform.googleapis.com/prediction/online/token_count
|
||||||
|
|
||||||
|
## 5. RAG Engine Cost Controls
|
||||||
|
|
||||||
|
- Delete stale documents from corpus
|
||||||
|
- Batch embedding for initial ingestion
|
||||||
|
- top_k=5 (not higher)
|
||||||
|
- Archive cold corpora to Cloud Storage lower tier
|
||||||
|
|
||||||
|
## 6. Memory Bank Cost Controls
|
||||||
|
|
||||||
|
- TTL: 7 days for short-lived session memories
|
||||||
|
- Firestore-backed for most use cases
|
||||||
|
- Index only user_id and session_id
|
||||||
|
|
||||||
|
## 7. Networking Cost Controls
|
||||||
|
|
||||||
|
- Private Google Access to avoid Vertex AI egress charges
|
||||||
|
- VPC routing for all agent traffic
|
||||||
|
- Cloud NAT only where unavoidable
|
||||||
|
|
||||||
|
## 8. Storage Lifecycle
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"rule": [{
|
||||||
|
"action": { "type": "Delete" },
|
||||||
|
"condition": { "age": 7 }
|
||||||
|
}]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
Apply with: `gcloud storage buckets update gs://BUCKET --lifecycle-file=lifecycle.json`
|
||||||
|
|
||||||
|
## 9. End-of-Day Teardown Checklist
|
||||||
|
|
||||||
|
- [ ] Run 03-teardown.sh
|
||||||
|
- [ ] Verify no active Agent Runtimes: `gcloud ai agent-engines list --region=us-central1`
|
||||||
|
- [ ] Verify no active Vertex AI endpoints: `gcloud ai endpoints list --region=us-central1`
|
||||||
|
- [ ] Check billing dashboard for anomalies
|
||||||
129
docs/GCP_Best_Practices.md
Normal file
129
docs/GCP_Best_Practices.md
Normal file
|
|
@ -0,0 +1,129 @@
|
||||||
|
# GCP Best Practices for Multi-Agent AI Systems
|
||||||
|
**OSVauco-NMTMD-GCOS | Last Updated: 2026-05**
|
||||||
|
|
||||||
|
## 1. Multi-Agent Architecture Patterns
|
||||||
|
|
||||||
|
### 1.1 Coordinator + Sub-Agent Pattern (ADK)
|
||||||
|
Use a root orchestrator agent that delegates to specialized sub-agents.
|
||||||
|
|
||||||
|
```python
|
||||||
|
root_agent = Agent(
|
||||||
|
name="orchestrator",
|
||||||
|
model="gemini-2.0-flash",
|
||||||
|
sub_agents=[greeting_agent, rag_agent, memory_agent, farewell_agent],
|
||||||
|
instruction="Delegate tasks to the appropriate sub-agent.",
|
||||||
|
before_model_callback=before_model_callback,
|
||||||
|
before_tool_callback=before_tool_callback,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
### 1.2 Graph Workflow Pattern
|
||||||
|
- SequentialAgent: ordered pipelines
|
||||||
|
- ParallelAgent: concurrent sub-tasks
|
||||||
|
- LoopAgent: retry/polling patterns
|
||||||
|
|
||||||
|
### 1.3 Session State and Memory
|
||||||
|
- Use session.state for short-lived conversation context
|
||||||
|
- Use Vertex AI Memory Bank for persistent long-term memory
|
||||||
|
- Tag memories with user_id and session_id for retrieval scoping
|
||||||
|
|
||||||
|
## 2. Model Selection
|
||||||
|
|
||||||
|
| Use Case | Recommended Model |
|
||||||
|
|----------|------------------|
|
||||||
|
| Complex reasoning / orchestration | gemini-2.0-flash or gemini-2.5-pro |
|
||||||
|
| Fast sub-agent calls | gemini-2.0-flash |
|
||||||
|
| Non-Google models | LiteLLM via LiteLlmModel |
|
||||||
|
| Embedding generation | text-embedding-005 |
|
||||||
|
| Code generation | gemini-2.5-pro |
|
||||||
|
|
||||||
|
## 3. Tool Design
|
||||||
|
|
||||||
|
- Single-responsibility per tool function
|
||||||
|
- Full docstrings — ADK uses these for LLM instructions
|
||||||
|
- Validate inputs inside tools; never trust LLM output blindly
|
||||||
|
- Use before_tool_callback for argument guardrails
|
||||||
|
- Return structured dicts, not raw strings
|
||||||
|
|
||||||
|
## 4. Vertex AI Integration
|
||||||
|
|
||||||
|
### 4.1 RAG Engine
|
||||||
|
**CONFIRMED** (source: Vertex AI RAG Engine SDK, google/adk-python samples, 2026-05):
|
||||||
|
- Use RagCorpus for document grounding
|
||||||
|
- Embedding model: text-embedding-005
|
||||||
|
- Chunking: 512 tokens, 50-token overlap
|
||||||
|
- Retrieval: top_k=5, vector_distance_threshold=0.7
|
||||||
|
|
||||||
|
**UNKNOWN / UNDOCUMENTED** (as of 2026-05):
|
||||||
|
- Full regional availability (eu-west3 issues known)
|
||||||
|
- Whether rag.import_files() is idempotent for same GCS path
|
||||||
|
- See: docs/Doc_Gaps_and_Open_Questions.md OQ-02
|
||||||
|
|
||||||
|
### 4.2 Memory Bank
|
||||||
|
**CONFIRMED** (source: Gemini Enterprise Agent Platform Memory Bank setup, 2026-05):
|
||||||
|
- load_memory_tool + preload_memory_tool from google.adk.tools
|
||||||
|
- Required role: roles/aiplatform.user
|
||||||
|
- SDK: google-cloud-aiplatform >= 1.111.0
|
||||||
|
|
||||||
|
**UNKNOWN / UNDOCUMENTED** (as of 2026-05):
|
||||||
|
- gcloud CLI equivalent for client.agent_engines.create()
|
||||||
|
- Whether tools auto-discover Memory Bank instance via project/region
|
||||||
|
- See: docs/Doc_Gaps_and_Open_Questions.md OQ-01, OQ-06
|
||||||
|
|
||||||
|
### 4.3 Security Controls Matrix
|
||||||
|
|
||||||
|
| Component | Data Residency | CMEK | VPC-SC | Access Transparency |
|
||||||
|
|-----------|---------------|------|--------|-------------------|
|
||||||
|
| Agent Platform | YES | YES | YES | YES |
|
||||||
|
| RAG Engine | NO | NO | YES | NO |
|
||||||
|
| Vector Search | YES | NO | YES | NO |
|
||||||
|
|
||||||
|
## 5. Observability
|
||||||
|
|
||||||
|
- Cloud Trace: end-to-end request tracing
|
||||||
|
- Cloud Logging: structured JSON logs with session_id; never log PII
|
||||||
|
- Cloud Monitoring: latency p50/p95/p99, error rates
|
||||||
|
- Error Reporting: auto-alert on new exception types
|
||||||
|
|
||||||
|
## 6. Deployment
|
||||||
|
|
||||||
|
- Cloud Run services (serverless, auto-scaling)
|
||||||
|
- Artifact Registry for container images
|
||||||
|
- min-instances=1 for latency-sensitive agents
|
||||||
|
- Cloud Build for CI/CD; --no-traffic deploy + HITL traffic switch
|
||||||
|
- Label all resources: env, team, cost-center, agent
|
||||||
|
|
||||||
|
## 7. MCP Integration
|
||||||
|
|
||||||
|
```python
|
||||||
|
from google.adk.tools.mcp_tool.mcp_toolset import MCPToolset, StdioServerParameters
|
||||||
|
|
||||||
|
tools, exit_stack = await MCPToolset.from_server(
|
||||||
|
connection_params=StdioServerParameters(
|
||||||
|
command="npx",
|
||||||
|
args=["-y", "@modelcontextprotocol/server-filesystem", "/tmp"]
|
||||||
|
)
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
## 8. A2A Protocol
|
||||||
|
|
||||||
|
- Agents expose AgentCard at /.well-known/agent.json
|
||||||
|
- A2A enables multi-vendor agent interoperability
|
||||||
|
- Secure A2A calls with service account tokens + VPC-SC
|
||||||
|
|
||||||
|
## 9. Handling Missing or Incomplete GCP Documentation
|
||||||
|
|
||||||
|
*Implements OSVauco-NMTMD-GCOS section 1.5.*
|
||||||
|
|
||||||
|
| Situation | Action |
|
||||||
|
|-----------|--------|
|
||||||
|
| API confirmed in GA docs or stable samples | Use it; cite source |
|
||||||
|
| API only in preview docs / blog post | Mark PREVIEW — validate before prod |
|
||||||
|
| API not found in any authoritative source | Omit; add to Open Questions |
|
||||||
|
|
||||||
|
## References
|
||||||
|
- ADK: https://google.github.io/adk-docs/
|
||||||
|
- Vertex AI RAG: https://cloud.google.com/vertex-ai/generative-ai/docs/rag-overview
|
||||||
|
- Gemini Enterprise Agent Platform: https://cloud.google.com/gemini-enterprise-agent-platform
|
||||||
|
- A2A: https://google.github.io/A2A/
|
||||||
Loading…
Reference in New Issue
Block a user