Agent-Relay v2

From SQLite to Kubernetes: Building and Deploying an Agent Relay Service
Building a small API is relatively easy.
Building one that can reliably register agents, deliver tasks, recover work after worker failures, persist state in PostgreSQL, run in containers, deploy to Kubernetes, and verify its own deployment is a very different engineering exercise.
This project, Agent Relay, started as a small self-contained SQLite implementation and evolved through six development stages into a containerized and Kubernetes-deployed service with automated testing and CI/CD.
The goal was not simply to make the application work.
The goal was to understand what happens when a small service moves from local development toward a real deployment workflow.
What is Agent Relay?
Agent Relay is a small FastAPI service responsible for coordinating agents and tasks.
At a high level:
Client / API / Dashboard
│
│ HTTP
▼
Agent Relay
FastAPI
│
▼
PostgreSQL
┌──────┼────────┐
│ │ │
Tasks Claims Agents
│
▼
Worker
Claim → Execute
→ Complete
The system supports agent registration, task creation and delivery, worker execution, task claims, lease expiration, recovery, and result submission.
The original implementation used SQLite as a self-contained local database. The persistence layer was later extended to support PostgreSQL for the containerized deployment.
From a Local Application to a Deployable System
One of the most useful aspects of this project was its incremental evolution.
Rather than starting with Kubernetes immediately, the system developed through a sequence of increasingly realistic engineering requirements.
Q1
│
├── Agents
├── Tasks
├── Claims
└── Results
│
▼
Q2
│
├── Worker execution
└── Automated tests
│
▼
Q3
│
├── Docker
└── Containerized FastAPI
│
▼
Q4
│
├── PostgreSQL
└── Persistent database
│
▼
Q5
│
├── Kubernetes
├── Kind
└── Persistent storage
│
▼
Q6
│
├── GitHub Actions
├── Test-gated build
├── SHA-based image tagging
└── Deployment verification
This progression made it possible to isolate and understand each engineering concern before adding the next layer.
1. Agent Registration and Authentication
Agents register through the API and receive an authentication token.
For example:
curl -X POST http://127.0.0.1:8000/api/v1/agents \
-H 'content-type: application/json' \
-d '{"name":"alice"}'
The returned token is subsequently used as a bearer token for authenticated operations.
For shared installations, registration can additionally be protected using an enrollment secret.
This creates a simple authentication boundary between the relay service and its workers.
2. Task Delivery and Worker Execution
The included deterministic worker performs a simple transformation:
input.upper()
For example:
Input:
hello agent relay
Output:
HELLO AGENT RELAY
Although the transformation itself is intentionally simple, the important part is the lifecycle surrounding it:
Create Task
│
▼
Claim Task
│
▼
Execute
│
▼
Complete
This provides a useful foundation for studying distributed task-delivery semantics without introducing unnecessary application complexity.
3. At-Least-Once Delivery
One of the more interesting parts of the system is its delivery model.
Agent Relay uses at-least-once task delivery.
A worker does not simply take a task permanently.
Instead, a task is claimed with a lease.
The default lease duration is:
60 seconds
Long-running workers send heartbeats while processing tasks.
If a worker disappears before completing the task, the claim eventually expires.
Another worker can then recover and claim the task.
Worker A
│
├── Claim
│
├── Heartbeat
│
└── Worker disappears
│
▼
Lease expires
│
▼
Worker B
│
└── Recover task
This makes failure recovery an explicit part of the application rather than something left entirely to the infrastructure layer.
4. Idempotent Completion
Terminal task operations also have an important safety property.
A terminal completion or failure requires:
the recipient's bearer token
the claim token
Repeated requests using the exact same claim token and result are treated idempotently.
A stale claim token, or a different result for an already-terminal claim, returns:
409 Conflict
This prevents workers from accidentally overwriting the final state of a task after ownership has changed.
5. PostgreSQL Persistence
The original SQLite implementation was intentionally kept simple.
SQLite is useful for a self-contained local starter because there is no external database dependency.
However, the containerized deployment uses PostgreSQL.
The application receives database configuration through environment variables instead of hard-coding credentials.
The Kubernetes deployment also separates database credentials into a Kubernetes Secret and uses persistent storage for PostgreSQL.
The resulting architecture is:
Kubernetes
│
├── Agent Relay Deployment
│ └── Agent Relay container
│
├── Agent Relay Service
│
├── PostgreSQL Deployment
│ └── PostgreSQL container
│
├── PostgreSQL Service
│
├── PostgreSQL Secret
│
└── PostgreSQL PVC
This separation allows the application and database to evolve independently while keeping credentials outside the application image.
6. Docker
The application was containerized using Docker.
The image can be built with:
docker build -t agent-relay:local .
The resulting image was successfully built as:
agent-relay:local
and was approximately 361 MB.
The container exposes port 8000:
docker run --rm \
-p 8000:8000 \
agent-relay:local
The same application artifact can then be loaded into the local Kubernetes cluster.
This creates an important deployment principle:
Application
│
▼
Docker Image
│
├── Local container
│
└── Kubernetes
The deployment artifact remains consistent between these environments.
7. Automated Testing
The project includes automated tests covering the main Agent Relay lifecycle.
Running:
uv run pytest -q
produces:
5 passed
The tests cover areas including:
agent authentication
sender/recipient access boundaries
task claiming
claim-token hashing
idempotent terminal requests
concurrent claims
lease expiration
task recovery
pagination
API error responses
dashboard asset serving
The test suite therefore validates more than basic HTTP availability.
It exercises important properties of the task-delivery protocol itself.
8. Kubernetes with Kind
After containerization and PostgreSQL persistence were working, the application was deployed to a local Kubernetes cluster using Kind.
The deployment contains:
Agent Relay Deployment
Agent Relay Service
PostgreSQL Deployment
PostgreSQL Service
PostgreSQL Secret
PostgreSQL PVC
The local Docker image is loaded into Kind:
kind load docker-image agent-relay:local --name kind
Then the Kubernetes manifests are applied:
kubectl apply -f k8s/
Deployment status can be checked with:
kubectl get pods
kubectl get deployments
kubectl get services
kubectl get pvc
And the application rollout can be verified using:
kubectl rollout status deployment/agent-relay
9. Health vs Readiness
A small but important design decision was separating liveness from readiness.
The service provides:
/health
/ready
/health verifies that the application process is alive.
/ready goes further by verifying database connectivity and schema availability.
This distinction matters in container orchestration.
An application process can be running while its database is unavailable.
Therefore:
Process alive
≠
Application ready
This is a small implementation detail, but it becomes increasingly important as applications move into orchestrated environments.
10. Kubernetes Port Forwarding
The local Kubernetes deployment can be accessed with:
kubectl port-forward service/agent-relay 18000:8000
The dashboard is then available at:
http://127.0.0.1:18000/
The port 18000 was intentionally chosen for the verification workflow to avoid collisions with unrelated applications using port 8000.
The deployment was successfully verified through the dashboard, including the application version:
Agent Relay v2
11. CI/CD with GitHub Actions
The final stage connected the application lifecycle to GitHub Actions.
The pipeline follows:
Git Push
│
▼
Run Tests
│
│ success
▼
Build Docker Image
│
▼
Tag Image with Git SHA
│
▼
Ensure Kind Cluster
│
▼
Export Kubeconfig
│
▼
Load Image into Kind
│
▼
Apply Kubernetes Manifests
│
▼
Update Deployment
│
▼
Wait for Rollout
│
▼
Port Forward
│
▼
Verify Dashboard
A particularly useful design choice is that deployment is test-gated.
pytest
│
├── failed ──► stop
│
└── passed
│
▼
Docker build
│
▼
Kubernetes deploy
A failed test therefore prevents the deployment stage from proceeding.
12. Git SHA-Based Image Tags
The Docker image is tagged using the Git commit SHA:
IMAGE_TAG: ${{ github.sha }}
This creates an immutable relationship between:
Git Commit
│
▼
Docker Image
│
▼
Kubernetes Deployment
Instead of relying only on a mutable tag such as:
latest
the deployed artifact can be traced back to the exact source revision that produced it.
For deployment debugging and reproducibility, this is a useful property.
13. CI/CD Deployment Verification
The pipeline does not stop after checking that a Kubernetes Pod exists.
After deployment, it performs an application-level verification through the dashboard.
The workflow:
kubectl apply -f k8s/
updates the Kubernetes resources, waits for the rollout, and then accesses the application through port forwarding.
The verification checks for:
Agent Relay v2
This creates a stronger deployment signal:
Container running
+
Kubernetes rollout successful
+
Application responding
+
Dashboard verification
14. Failure Recovery
The leased task model also provides a practical failure-recovery mechanism.
A worker can be deliberately slowed:
uv run python main.py worker \
--credentials ./uppercase-credentials.json \
--slow-seconds 75 \
--worker-id slow-laptop
If the worker terminates during execution, the claim remains leased until expiration.
After the lease expires:
Worker A
│
▼
Claim Task
│
▼
Worker disappears
│
▼
Lease expires
│
▼
Worker B
│
▼
New claim
│
▼
Attempt count increases
This demonstrates that recovery is part of the application protocol.
Engineering Lessons
Several engineering ideas became particularly clear through this project.
1. Start simple, then add operational complexity
The project did not require Kubernetes on day one.
The evolution from:
SQLite
to:
PostgreSQL
then:
Docker
then:
Kubernetes
and finally:
CI/CD
made each new operational concern easier to reason about.
2. Separate protocol from persistence
The storage layer was kept separate from the HTTP protocol.
Conceptually:
HTTP / API
│
▼
Storage Operations
│
├── SQLite
│
└── PostgreSQL
The original SQLite implementation uses an explicit BEGIN IMMEDIATE transaction strategy because SQLite does not provide PostgreSQL's FOR UPDATE SKIP LOCKED.
This separation made it possible to extend persistence without changing the core Agent Relay lifecycle.
3. Deployment verification should test the application
A Kubernetes rollout succeeding does not necessarily mean the application is usable.
That is why the CI/CD workflow continues beyond:
kubectl rollout status
and performs an application-level dashboard verification.
4. Failure handling belongs in the design
Worker failure, lease expiration, retries, and idempotent terminal operations are not merely infrastructure concerns.
They are part of the application protocol.
Final Architecture
The completed development-to-deployment workflow is:
Developer
│
▼
Git Repository
│
▼
GitHub Actions
│
├── pytest
│
├── Docker build
│
└── Kubernetes deployment
│
▼
Kind Cluster
│
┌────┴────┐
▼ ▼
Agent Relay PostgreSQL
│
▼
Dashboard
The project successfully demonstrates:
agent registration and authentication
task creation and delivery
worker execution
at-least-once delivery
lease and heartbeat handling
task recovery
idempotent terminal requests
automated testing
PostgreSQL persistence
Docker containerization
Kubernetes deployment
persistent PostgreSQL storage
rollout verification
dashboard verification
GitHub Actions CI/CD
test-gated deployment
Git SHA-based image tagging
end-to-end deployment verification
Conclusion
Agent Relay started as a small local protocol implementation.
It ended as a practical demonstration of the complete path from application code to a running Kubernetes service:
Code
↓
Tests
↓
Docker
↓
Git SHA Image
↓
Kind
↓
Kubernetes
↓
PostgreSQL
↓
Agent Relay
↓
Dashboard / API
The most valuable part of the project was not the individual technologies.
It was understanding how the pieces connect.
Application design → persistence → containerization → orchestration → CI/CD → deployment verification.
That progression provides a practical foundation for building and operating more complex distributed and AI-enabled systems in the future.
GitHub repo: https://github.com/ketut-garjita/agent-relay
