Implemented a playbook-driven external smoke test framework and initial scenarios with DNS-inclusive checks.
What was added/changed
New files:
smoke/config.py – playbook schema, env overrides, loader, scenario filtering
smoke/runner.py – runner with three scenarios, waiters, DNS/TCP/HTTP probes, JSON reporting
smoke/playbook.example.yaml – sample playbook using IP-based defaults with DNS checks
smoke/README.md – usage and operation docs
smoke/init.py – package marker
API client uplift:
Added IaaSClient.container_lifecycle() to call POST /workloads/containers/<id>/lifecycle/<action> for future restart tests
Leveraged existing client methods:
IaaSClient.get_container() and IaaSClient.delete_container()
IaaSClient.get_container_pods() and IaaSClient.get_container_pod()
Server routes/flow relied on:
Create (async, DB-first): /workloads/containers
Worker status update: update_container_workload()
Pod introspection with DNS and port_forwardings: get_pods(), get_pod()
Allocation + DNS/Tunnel creation: allocate_and_dispatch(), allocate_ports_for_container(), CloudflareTunnelManager
Scenarios implemented (playbook-driven)
create-nginx
Creates a container in a new pod (nginx:stable-alpine by default), requests port 80 with use_dns: true.
Waits for allocation then running.
Discovers ingress from pod.port_forwardings (dns_record_hostname or external_ip_address:external_port).
Probes DNS resolution, TCP connectivity, and HTTP GET / expecting 200 and “nginx”.
Teardown deletes the container, then validates:
TCP connectivity to previously known ingress fails
Port forwarding entries for the container are removed from the pod
Code: _scenario_create_container()
add-second-container
Adds a container to the pod created above, waits running, and probes the new container.
Also probes original container still serves traffic.
Deletes only the newly added container.
Code: _scenario_add_to_pod()
negative-image
Uses a bogus image (thisdoesnotexist.invalid:never), expects final status launch_failed.
Validates no ingress is assigned.
Cleans up the failed container record.
Code: _scenario_negative_image()
Runner capabilities
DNS checks are enforced based on playbook defaults; uses DNS if available and also probes external IP:port as fallback unless dns_required is strictly enforced.
Robust waiters with polling and timeouts: allocation, running, deletion.
Data-plane probes:
DNS resolution: dns_resolve()
TCP connect: tcp_connect()
HTTP GET with redirects and content assertions: http_get()
JSON report with timings and per-endpoint assertions written to configurable path.
How to run
Set env vars or edit the playbook:
API_BASE_URL (default http://172.17.0.1:5000/api)
API_KEY
VDC_ID (target UUID)
DNS_REQUIRED=true (enforced in example playbook)
Optional: WEBSOCKET_SERVER_URL for documentation purposes
Review and update smoke/playbook.example.yaml.
Execute:
python smoke/runner.py --playbook smoke/playbook.example.yaml --report smoke/out/report.json
The runner prints summary and writes detailed JSON results.
Notes and extension points
Container restart validations can leverage IaaSClient.container_lifecycle() and assert data-plane blip/resume while worker transitions via container_lifecycle_action().
Additional coverage can be added for volumes, networks, VMs, and pod lifecycle actions reusing the same framework.
The framework stays external: only public API calls and real-world network probes; no internal DB calls.
This delivers the initial smoke test infrastructure and the required test coverage: create/delete container with DNS and data-plane verification, add container to an existing pod with verification for both, and a negative “dodgy image” test that must fail correctly.