> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tensormesh.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Architecture

Tensormesh Platform installs from a single Helm chart. The chart runs a controller (the
**operator**) that reconciles `LMCacheEngine` custom resources into a **DaemonSet of LMCache
cache servers**, stands up a fleet **coordinator**, and (when observability is enabled)
wires those engines to an **OpenTelemetry Collector** that fans metrics and traces out to
Prometheus and Tempo.

## What gets created, and by whom

Every piece follows the same pattern: **the chart writes a declaration (a CR); a controller
reconciles it into a running thing.** The chart runs exactly **one** controller itself — the
operator. The Collector and Tempo are reconciled by operators **you pre-install**;
the chart only submits their CRs. This table is the precise "which thing belongs where":

| Declaration (the chart creates this) | Reconciled by | Becomes | Scope |
| - | - | - | - |
| `LMCacheEngine` CR | **operator** (this chart) | engine DaemonSet (1 pod per GPU node) | Namespaced |
| `LMCacheCoordinator` CR | **operator** (this chart) | coordinator Deployment + ClusterIP Service (`:9300`) | Namespaced |
| `CacheBlendEngine` CR *(Preview)* | **operator** (this chart) | CacheBlend DaemonSet (1 pod per GPU node) | Namespaced |
| `OpenTelemetryCollector` CR | **OpenTelemetry Operator** (prerequisite) | OTel Collector pod | Namespaced |
| `TempoMonolithic` CR | **Tempo Operator** (prerequisite) | Tempo instance | Namespaced |
| `ServiceMonitor` | **Prometheus Operator** (prerequisite) | a scrape target (no new pod) | Namespaced |
| CRDs `lmcacheengines`, `lmcachecoordinators`, `cacheblendengines` (`.lmcache.lmcache.ai`) | None (registered with the API server) | define the `LMCacheEngine`, `LMCacheCoordinator`, and `CacheBlendEngine` kinds | **Cluster** |

The operator itself runs as a plain `Deployment` (no CR needed) and owns the
engine DaemonSet's lifecycle.

## Data flow

```text theme={null}
vLLM pod  ──hostIPC / MP connector──▶  LMCache engine
LMCache engine  ──OTLP gRPC───────────▶  OTel Collector
OTel Collector  ──Prometheus scrape───▶  Prometheus   (via ServiceMonitor)
OTel Collector  ──OTLP traces─────────▶  Tempo / external backend
```

| Hop | Transport | Port |
| - | - | - |
| vLLM → engine | shared-memory IPC (`hostIPC`) | — (same node) |
| engine → Collector | OTLP gRPC | `4317` |
| Collector → Prometheus | scrape of the Collector's exporter | `8889` |
| Collector → itself | `otelcol_*` internal metrics | `8888` |
| Collector → Tempo / external | OTLP | per backend |

See [Observability](/v1.0.0/observability) for enabling and verifying this path.

## Resources the chart creates

Everything the chart can render, with its **scope**. This is the table to consult before
a cleanup, because **cluster-scoped resources survive a namespace delete** and must be
removed explicitly.

| Resource | API group | Scope | Created when | Cleanup notes |
| - | - | - | - | - |
| **CustomResourceDefinitions** `lmcacheengines`, `lmcachecoordinators`, `cacheblendengines` (`.lmcache.lmcache.ai`) | `apiextensions.k8s.io` | **Cluster** | `crds.enabled` (default `true`) | ⚠️ All three **kept on `helm uninstall`** (`resource-policy: keep`). Survive namespace delete. Remove with `kubectl delete crd …`. |
| **ClusterRole** (operator) | `rbac.authorization.k8s.io` | **Cluster** | `operator.rbac.create` | ⚠️ Survives namespace delete; sweep manually. |
| **ClusterRoleBinding** (operator) | `rbac.authorization.k8s.io` | **Cluster** | `operator.rbac.create` | ⚠️ Survives namespace delete; sweep manually. |
| **ClusterRole / ClusterRoleBinding** (pre-delete hook) | `rbac.authorization.k8s.io` | **Cluster** | `helm uninstall` hook | Created and removed by the hook; can orphan if uninstall fails mid-way. |
| `LMCacheEngine` (CR) | `lmcache.lmcache.ai/v1alpha1` | Namespaced | `engine.enabled` (default `true`) | Has a `lmcache.ai/cleanup` **finalizer**: if the operator is gone, it blocks namespace deletion. See [troubleshooting](/v1.0.0/installation/troubleshooting#cleaning-up-an-inconsistent-or-orphaned-install). |
| `LMCacheCoordinator` (CR) | `lmcache.lmcache.ai/v1alpha1` | Namespaced | `coordinator.enabled` (default `true`) | Goes with the namespace. |
| `CacheBlendEngine` (CR) *(Preview)* | `lmcache.lmcache.ai/v1alpha1` | Namespaced | `cacheBlend.enabled` | Goes with the namespace. |
| `OpenTelemetryCollector` (CR) | `opentelemetry.io` | Namespaced | `observability.enabled` | Goes with the namespace. |
| `TempoMonolithic` (CR) | `tempo.grafana.com` | Namespaced | `observability.traces.tempoCR.enabled` | Goes with the namespace. |
| `ServiceMonitor` (CR) | `monitoring.coreos.com` | Namespaced | `operator.serviceMonitor.enabled` and/or `observability.prometheus.serviceMonitor.enabled` | Goes with the namespace. |
| Deployment (controller-manager) | `apps` | Namespaced | `operator.enabled` | Goes with the namespace. |
| DaemonSet (engine) | `apps` | Namespaced | created by the operator from the `LMCacheEngine` CR | Removed when the CR/namespace is. |
| Deployment + Service (coordinator) | `apps` / core | Namespaced | created by the operator from the `LMCacheCoordinator` CR | Removed when the CR/namespace is. |
| DaemonSet (CacheBlend engine) *(Preview)* | `apps` | Namespaced | created by the operator from the `CacheBlendEngine` CR | Removed when the CR/namespace is. |
| Service (metrics) | core | Namespaced | `operator.enabled` | Goes with the namespace. |
| ServiceAccount (operator) | core | Namespaced | `operator.serviceAccount.create` | Goes with the namespace. |
| Role / RoleBinding (operator) | `rbac.authorization.k8s.io` | Namespaced | `operator.rbac.create` | Goes with the namespace. |
| ServiceAccount + RoleBinding (privileged SCC) | core / `rbac…` | Namespaced | `openshift.enabled` | Goes with the namespace. Binds to the built-in `system:openshift:scc:privileged` ClusterRole (**do not delete that**). |
| NetworkPolicy | `networking.k8s.io` | Namespaced | `operator.networkPolicy.enabled` | Goes with the namespace. |
| Job + ServiceAccount + Role + RoleBinding (pre-delete hook) | batch / core / `rbac…` | Namespaced | `helm uninstall` hook | Transient; removed after the hook runs. |

### What a namespace delete leaves behind

If you `kubectl delete namespace` instead of `helm uninstall`, these **cluster-scoped**
resources orphan and need explicit cleanup:

1. **CRDs** `lmcacheengines`, `lmcachecoordinators`, and `cacheblendengines` (all `.lmcache.lmcache.ai`)
2. **ClusterRole** for the operator
3. **ClusterRoleBinding** for the operator

```bash theme={null}
kubectl delete crd lmcacheengines.lmcache.lmcache.ai lmcachecoordinators.lmcache.lmcache.ai cacheblendengines.lmcache.lmcache.ai
kubectl get clusterrole,clusterrolebinding | grep -i tensormesh   # delete what it lists
```

See [Troubleshooting → Cleaning up an inconsistent or orphaned install](/v1.0.0/installation/troubleshooting#cleaning-up-an-inconsistent-or-orphaned-install)
for the full recipe (including clearing the `LMCacheEngine` finalizer first).

## Next steps

<CardGroup cols={2}>
  <Card title="Configuration" icon="sliders" href="/v1.0.0/reference/configuration">
    Every `values.yaml` key and example overlays.
  </Card>

  <Card title="Observability" icon="chart-line" href="/v1.0.0/observability">
    The OTel Collector / Tempo / Prometheus wiring in detail.
  </Card>
</CardGroup>
