Process-liveness only — never touches a dependency, so a slow/unreachable
backend (especially the OFF-cluster LLM or the GPU OCR servers) can never fail
the liveness probe and needlessly restart the pod. Use /v1/health (deep) for
dashboards and /v1/health/ready for readiness.