From 7e85fc8e9201402c720d38ba5622e25b31bd20e5 Mon Sep 17 00:00:00 2001 From: arlando Date: Fri, 28 Aug 2026 16:57:18 -0400 Subject: [PATCH] PLAT-95: add Teleport access spike research brief Scaffold discovery doc for monitoring, multi-cluster/DB access, AWS/network constraints, and auditing. Links from access/ stubs to Jira spike PLAT-95. --- access/endpoints/stargate.md | 7 +- access/services/README.md | 2 + docs/research/teleport-access-spike.md | 156 ++++++++++++++++++ .../teleport-access-spike-ticket-draft.md | 47 ++++++ 4 files changed, 211 insertions(+), 1 deletion(-) create mode 100644 docs/research/teleport-access-spike.md create mode 100644 jira/plans/examples/teleport-access-spike-ticket-draft.md diff --git a/access/endpoints/stargate.md b/access/endpoints/stargate.md index bbb4ac1..ec73d3f 100644 --- a/access/endpoints/stargate.md +++ b/access/endpoints/stargate.md @@ -16,7 +16,8 @@ sources_to_synthesize: # Stargate (Teleport) **Public URL:** `https://stargate.odieplat.io` -**Jira:** Epic E01 → Story [PLAT-100](https://catalystsoftware.atlassian.net/browse/PLAT-100) (Platform Access — Stargate) +**Jira:** Epic E01 → Story [PLAT-10](https://catalystsoftware.atlassian.net/browse/PLAT-10) (Platform Access — Stargate) +**Research spike:** [PLAT-95](https://catalystsoftware.atlassian.net/browse/PLAT-95) — monitoring, DB access, AWS/network discovery ([brief](../../docs/research/teleport-access-spike.md)) ## Purpose @@ -33,6 +34,10 @@ kubectl get nodes Expected pattern: Google SSO once per day, then `tsh` session for kubectl and app access. +## Open research (spike) + +Teleport as a unified gateway for kubectl, DB access, and auditing across regions/VPCs is under active discovery. See [teleport-access-spike.md](../../docs/research/teleport-access-spike.md) and Jira spike PLAT-95. + ## Related - [ArgoCD](argocd.md) — same hub cluster, sequenced after Stargate MVP diff --git a/access/services/README.md b/access/services/README.md index 69b4e6e..c7b95d4 100644 --- a/access/services/README.md +++ b/access/services/README.md @@ -48,6 +48,8 @@ Platform-bots agents use **cluster-internal Service endpoints** and MCP between Future: register IKG MCP and other apps in Teleport when hub rollout completes. v1: `kubectl port-forward` only — see [mcps/interlink-map.md](../mcps/interlink-map.md). +**Research:** Multi-region DB access, monitoring, VPC/egress, and audit requirements — [teleport-access-spike.md](../../docs/research/teleport-access-spike.md) (Jira PLAT-95). + ```bash # PLACEHOLDER — Teleport app name TBD # tsh apps login diff --git a/docs/research/teleport-access-spike.md b/docs/research/teleport-access-spike.md new file mode 100644 index 0000000..3110351 --- /dev/null +++ b/docs/research/teleport-access-spike.md @@ -0,0 +1,156 @@ +--- +title: "Teleport access — monitoring, tooling, and multi-resource discovery" +tags: [eng-information, platform-bots, research] +last_updated: "2026-08-28" +status: draft +audience: [engineers, agents] +gaps: + - "Inventory of DB types, regions, and current access paths per environment" + - "Whether a centralized access tool exists beyond Teleport candidates" + - "VPC / security-group / NACL constraints for UDP/TCP from spokes to hub" + - "Teleport edition (OSS vs Enterprise) for DB protocol and audit features" +sources_to_synthesize: + - "argocd-tele: workspace/docs/eks-teleport-platform-research.md" + - "argocd-tele: workspace/docs/jira-space-organization.md (E01)" + - "platform-tools/access/endpoints/stargate.md" +jira: "PLAT-95" +--- + +# Teleport access spike — research brief + +**Jira:** [PLAT-95](https://catalystsoftware.atlassian.net/browse/PLAT-95) (Spike under Epic E01 → Story PLAT-10) +**Hub endpoint:** `https://stargate.odieplat.io` +**Canonical prior art:** `argocd-tele` → `workspace/docs/eks-teleport-platform-research.md` + +## Goal + +Discover what is required to use **Teleport (Stargate)** as the primary **human access gateway** for: + +- Kubernetes clusters (multi-region, private API endpoints) +- Databases (heterogeneous engines, regions, and VPC placement) +- Auditing and operational monitoring of access sessions + +This spike is **discovery only** — no production changes, no secrets in git, no personal machine paths. + +## Out of scope + +- ArgoCD hub migration (separate Story PLAT-108) +- Full Teleport hub deployment (tracked under PLAT-43 and subtasks) +- Populating `access/` with verified connection strings (see [PLAT-92](https://catalystsoftware.atlassian.net/browse/PLAT-92)) + +## Research questions + +### 1. Monitoring and tooling + +| Question | Notes | +|----------|-------| +| What should we monitor on the Teleport proxy and agents? | Health, auth failures, agent disconnects, cert expiry | +| Where do metrics/logs land today? | Hub `kube-prometheus-stack` pattern from Leviosa; audit to S3 | +| What alerts matter for access outages? | Agent cannot dial proxy; OIDC failures; NLB target unhealthy | +| CLI ergonomics | `tsh status`, `tsh kube ls`, `tsh db ls` — document expected engineer workflow | +| IDE / agent integration | Port-forward vs Teleport app access for MCP services (stub in `access/mcps/`) | + +**Deliverable:** Recommended observability stack for Teleport on `platform-eks` (metrics, logs, audit export, on-call runbook outline). + +### 2. AWS changes that may be required + +| Area | Hypothesis | Verify | +|------|------------|--------| +| **NLB** | Internet-facing TCP passthrough to Teleport `:443` (do not terminate TLS at LB) | Confirm with Teleport HA reference architecture | +| **VPC / subnets** | Reuse shared-services `10.200.0.0/16`; new EKS subnets | Subnet capacity, route tables | +| **Egress from spokes** | Kube agents dial **outbound** to public proxy | NACLs, NAT, firewall rules per spoke | +| **TGW / peering** | Required for ArgoCD hub → spoke API; **optional** for Teleport-only kubectl | Document which paths need hub-to-spoke vs spoke-to-hub | +| **VPC flow logs** | Parquet/Hive to S3 (required) | Wire shared-services workspace to existing flow-log module | +| **IAM / Pod Identity** | Hub uses EKS Pod Identity (not IRSA default) | ALB controller, ExternalDNS, ESO associations | +| **Secrets** | Google OIDC client secret via External Secrets | Secrets Manager → ESO pattern | +| **Audit storage** | Teleport session recordings + audit events → S3 | Bucket policy, retention, encryption | + +**Deliverable:** AWS change checklist (Terraform workspaces affected, net-new vs reuse). + +### 3. Kubernetes cluster access — context switching best practices + +| Topic | Research | +|-------|----------| +| **Private EKS APIs** | All spokes use `endpoint_public_access = false` — Teleport kube agents inside each cluster | +| **Registration model** | Hub registers clusters; engineers `tsh kube login ` | +| **Context hygiene** | Avoid committing kubeconfig; use `tsh` cert-based contexts; document `~/.config//config.yaml` pattern only | +| **Multi-cluster UX** | Naming convention for clusters; dev vs prod RBAC via Teleport roles | +| **Break-glass** | SSM bastion pattern (`eks-connect.sh`) until Teleport spoke rollout complete | +| **Rollout order** | MVP hub → `leviosa-dev-eks` pilot → broader spokes (PLAT-53) | + +**Deliverable:** Recommended engineer workflow doc section for `access/endpoints/stargate.md` (placeholders only until verified). + +### 4. Database access — multi-region, multi-engine + +| Challenge | Discovery needed | +|-----------|------------------| +| **Engine diversity** | Postgres, MySQL, Redis, DynamoDB, etc. — which support Teleport DB access vs need other paths | +| **Regional spread** | DBs in different regions/clusters/VPCs — agent placement per region | +| **Existing patterns** | VPN, SSM tunnels, PrivateLink, Pomerium — what is used today per DB class | +| **Teleport DB access** | `tsh proxy db` / `tsh db connect` — protocol support, TLS requirements | +| **Centralized gateway** | Is Teleport the single pane, or hybrid (Teleport for some, PrivateLink for service-to-service)? | +| **Customer data** | Session recordings and query audit — minimize/obfuscate customer identifiers in logs | + +**Deliverable:** DB access matrix (engine × region × current path × Teleport feasibility × blockers). + +### 5. Network constraints — VPC, UDP, TCP + +| Path | Protocol | Discovery | +|------|----------|-----------| +| Engineer → Stargate | TCP 443 (HTTPS / Teleport TLS) | Public NLB; optional corp IP allowlist | +| Spoke agent → Stargate | TCP 443 outbound | Verify egress per spoke VPC | +| `tsh proxy db` | TCP local forward | Client-side; document ports | +| Teleport tunnel / reverse tunnel | TCP (multiplexed) | Confirm no UDP requirement for MVP | +| Legacy VPN (OpenVPN) | UDP/TCP | Document overlap with Teleport rollout; retirement path (RFW) | +| ArgoCD → spoke API | TCP 443 | Separate fabric requirement (TGW) | + +**Deliverable:** Network requirements table + list of VPCs where egress to `stargate.odieplat.io:443` is blocked or unknown. + +### 6. Auditing and compliance + +| Capability | Question | +|------------|----------| +| **Session recording** | K8s exec/SSH — required? Storage and retention policy | +| **DB query audit** | Teleport DB audit events — sufficient for compliance asks? | +| **Identity** | Google OIDC email allowlist → Teleport roles (no group dependency) | +| **Revocation** | `tctl lock` / session termination — runbook | +| **Export** | SIEM integration (if any); correlation with VPC flow logs | +| **Privacy** | No customer names/tenant IDs in audit exports committed to git | + +**Deliverable:** Audit requirements doc + gap list vs Teleport OSS vs Enterprise. + +## Suggested investigation order + +1. Read `eks-teleport-platform-research.md` executive summary and open questions. +2. Inventory current human access paths (VPN, SSM, Pomerium) — **names only**, no hostnames with customer data. +3. Confirm spoke egress to `stargate.odieplat.io:443` (PLAT-54 pattern). +4. Draft DB matrix with platform/SRE input. +5. Propose monitoring + audit architecture for hub MVP. +6. File follow-up Tasks/Stories from findings (do not bulk-move existing tickets). + +## Acceptance criteria (spike) + +- [ ] Written findings added to this doc or linked ADR (privacy-safe) +- [ ] DB access matrix drafted (engines × regions × path × Teleport fit) +- [ ] AWS change checklist with Terraform workspace owners +- [ ] Network/egress gap list per spoke class (dev/prod/EU) +- [ ] Monitoring and audit recommendations for Teleport on `platform-eks` +- [ ] Follow-up Tasks created in Jira with clear owners (max 5 — no bulk ticket moves) +- [ ] `access/endpoints/stargate.md` updated with spike link only (no secrets) + +## Related tickets + +| Key | Summary | +|-----|---------| +| PLAT-4 | Epic E01 — Platform hub — Teleport & ArgoCD | +| PLAT-10 | Story — Platform Access — Stargate (Teleport) | +| PLAT-43 | Task — Teleport hub at stargate.odieplat.io | +| PLAT-53 | Task — Dev spoke pilot (leviosa-dev-eks) | +| PLAT-92 | Synthesize verified access patterns into `access/` | + +## References + +- [Stargate endpoint stub](../../access/endpoints/stargate.md) +- [MCP interlink map](../../access/mcps/interlink-map.md) +- [Security policy](../security.md) +- [How we work — privacy](../how-we-work.md#privacy) diff --git a/jira/plans/examples/teleport-access-spike-ticket-draft.md b/jira/plans/examples/teleport-access-spike-ticket-draft.md new file mode 100644 index 0000000..aedda0f --- /dev/null +++ b/jira/plans/examples/teleport-access-spike-ticket-draft.md @@ -0,0 +1,47 @@ +# Teleport access spike — PLAT-95 + +**Key:** [PLAT-95](https://catalystsoftware.atlassian.net/browse/PLAT-95) +**Type:** Spike +**Parent:** PLAT-4 (Epic E01 — Platform hub — Teleport & ArgoCD) +**Related story:** PLAT-10 (Platform Access — Stargate) +**Labels:** `stargate`, `teleport`, `eng-information`, `spike`, `access`, `tier:2` + +## Summary + +Spike: Teleport monitoring, multi-cluster/DB access, AWS and network discovery + +## Description + +Investigate how **Teleport (Stargate)** at `https://stargate.odieplat.io` should serve as the human access gateway for Kubernetes and databases across regions, clusters, and VPCs. + +### Scope + +**In scope (discovery):** + +1. **Monitoring and tooling** — metrics, logs, alerts, and engineer CLI workflow (`tsh`, context switching) +2. **AWS changes** — NLB, VPC/subnets, egress, flow logs, Pod Identity, audit S3, Secrets Manager +3. **Kubernetes access** — best practices for private EKS APIs, kube agents, RBAC, rollout order +4. **Database access** — heterogeneous engines and regions; Teleport DB vs VPN/PrivateLink/SSM today +5. **Network constraints** — VPC requirements, TCP/UDP paths, spoke egress to public proxy +6. **Auditing** — session recording, DB audit, identity/revocation, privacy (no customer data in exports) + +**Out of scope:** ArgoCD hub (PLAT-108), Teleport hub deployment tasks (PLAT-43+), populating `access/` with verified secrets (PLAT-92). + +### Prior art + +- `argocd-tele` → `workspace/docs/eks-teleport-platform-research.md` +- `platform-tools` → `docs/research/teleport-access-spike.md` + +### Deliverables + +- DB access matrix (engine × region × current path × Teleport feasibility) +- AWS change checklist with workspace owners +- Network/egress gap list per spoke class +- Monitoring + audit recommendations for Teleport on `platform-eks` +- Up to 5 follow-up Tasks with owners (no bulk ticket moves) + +### Rules + +- No secrets, tokens, or personal machine paths in findings committed to git +- Obfuscate/minimize customer names and tenant identifiers in any examples +- Use `~/.config//config.yaml` for local config references only