move multiprocessing manager initialization from constructor to run() method in both DockerMonitor and LibvirtMonitor to avoid issues with forking managers across processes
fix(worker): prevent stop event errors during shutdown
add null checks for stop_event and manager before attempting to use or shutdown them in monitor classes
refactor(worker): simplify process management in main
replace processes list with direct worker_proc handling for cleaner shutdown logic
feat(worker): add error handling for monitor process startup
wrap monitor process creation in try-except blocks to catch and log startup failures
- Move DockerMonitor and LibvirtMonitor instantiation into the worker process after successful join
- Add support for disabling monitors via no_docker and no_libvirt flags
- Implement automatic restart of worker process on unexpected termination
- Add reconnection logic to WorkerClient with exponential backoff
- Ensure container reconciliation runs on both initial join and reconnections
Introduce handling for pods with assigned but offline hosts in the retry task.
Increment allocation attempts, apply exponential backoff with jitter, and
schedule next retry accordingly. Skip immediate reallocation if host is offline
to prevent unnecessary attempts. Configuration values are used for base delay,
backoff factor, jitter, and max delay.
Refactored the container allocation logic to improve migration support:
- Moved NSController handling directly into main allocation flow
- Added migration-aware port allocation with existing port reuse
- Simplified function signatures by removing ns_lp parameter dependency
- Moved DB commit earlier in the process for better transaction management
- Improved stale port mapping cleanup with proper error handling
- Fixed typos in logging messages
- Add send_full_statuses_for_host to emit a full snapshot after
reconciliation for convergence. Supports "host" or "pod" scope via
WORKER_FULL_STATUS_SCOPE, filters by managed_by label, and skips
neutral states.
- Map Docker container statuses to websocket_server event types in
WorkerClient: "running" -> running, "exited"/"dead"/"stopped" -> die,
skip others. Align details.status and Action accordingly to avoid
unsupported events.
- Abort launching regular containers if NSController launch fails and
improve logging/injection behavior.
These changes improve state convergence and robustness while preventing
invalid event emissions.
- Update host assignment condition to reassign if workload host ID is set but does not match the selected host
- Add debug logging for host selection and pre-populated host cases
- Add TODO for setting pod state to allocated after container allocation
Introduce `ENABLE_TASK_ASSIGNMENT_DEBUG` configuration to control verbose
logging in task assignment, Redis subscription, and worker management flows.
This reduces log noise by default while allowing detailed debugging when
enabled.
Integrate OVSBridgeScannerTask into WorkerClient to scan for local OVS
bridges using ovs-vsctl, report them to the API endpoint, and handle
initialization with error logging. Includes retry logic for API calls
and graceful fallback if OVS is unavailable.
Add GET endpoint to fetch non-deleted OVS bridge names for a workload host.
Add POST endpoint to update OVS bridges: validate input, soft-delete removed
bridges, add new bridges, update timestamps for unchanged ones, with error
handling and logging.
Introduce support for container states (present, absent, started, stopped,
restarted) via xcloudify_containers_state variable in Ansible role.
Update container specs and module to use 'cpu' instead of 'cpu_shares' for
allocation, with validation and API adjustments.
Update examples and tasks for state handling, debug output, and error conditions.
BREAKING CHANGE: 'cpu_shares' renamed to 'cpu' in container configurations and module options.
Refactor HasSufficientResources to support arbitrary resource types via
host.get_resource_utilization(), applying per-resource overcommit ratios
(default 1.0). Introduce dynamic aggregation of pod resource requirements
from WorkloadResourceUsage database entries, replacing hardcoded minima.
Integrate the filter early in select_host_for_pod pipeline when CPU/RAM
requirements exceed zero, preserving filter order for efficiency.
Adjust global cloudflared sidecar limits: CPU to 1 (from 3), RAM to 64MiB
(from 132MiB) for reduced overhead.
- Introduce CloudflareReconciliationWorker to reconcile Cloudflare
tunnels/DNS with DB and clean up stale resources
- Add Celery task tasks.cloudflare_reconciliation and schedule it
every 5 minutes via beat
- Provide standalone script to run the reconciliation manually
- Add optional scheduler helper for periodic runs
Also:
- Update tunnel lookup to not ignore tunnels in down status so
cleanup can find and delete them
No breaking changes.
- modify `cleanup_cloudflare_resources` to accept workload object instead of container id
- remove redundant workload lookup inside the function
- update usage in `process_container_deletion` to pass container object
- adjust resource usage assignments in `ensure_tunnel_and_dns` to use sidecar id
- remove obsolete `_cleanup_nscontroller_cloudflare_resources` function
refactor resource limit definitions by using global variables for cpu and memory, and track resource usage in WorkloadResourceUsage entries for the NSController workload
Move all workload and pod deletion logic into a new unified deletion handler module. This includes marking workloads as pending-deleted, cleaning up associated resources (port forwardings, volume mappings, network ports, Cloudflare DNS records and tunnels), and updating pod statuses after container deletions. The change reduces code duplication across multiple files and ensures consistent cleanup behavior.
The new app/utils/deletion_handler.py module contains functions for:
- mark_workload_pending_deleted: marks workloads and associated resources as pending-deleted
- cleanup_workload_resources: performs actual resource deletion
- cleanup_cloudflare_resources_for_workload: handles Cloudflare-specific cleanup
- mark_pod_pending_deleted: marks all containers in a pod for deletion
- update_pod_status_after_deletion: updates pod status based on container states
Modified existing code to use these new consolidated functions instead of repeating similar logic in multiple places.
Move NoFreePortsError to dedicated exceptions module and add CloudFlareException for
Cloudflare-specific failures. Update ensure_tunnel_and_dns to raise CloudFlareException
on errors. Integrate set_waiting_host_and_next_retry in allocation_retry and
process_workload_request tasks for consistent failure handling with exponential backoff.
Handle pod ID as string in set_waiting_host_and_next_retry and add logging.
Add get_resource_utilization to WorkloadHost to compute available and used
resources dynamically from pooled resources and active workloads.
Add get_resource_usage to Workload to return list of used resources for the
workload.
Fix minor typo in TODO comment and add SQLAlchemy func import.
Renamed the CPU resource parameter from 'cpu_shares' to 'cpu' across API
validation, payload building, resource usage tracking, and worker tasks for
consistency and simplicity. Adjusted default mem_limit from 128MB to 129MB.
Introduced global constants for NSController resources (5 CPU, 199MB) and set
tunnel sidecar limits (3 CPU, 257MB). Removed unused import.
BREAKING CHANGE: Container API payloads now expect 'cpu' key instead of 'cpu_shares
Introduce _enrich_host_data helper to integrate resource utilization data
into host responses, replacing manual resource queries.
- Update get_workload_host and get_workload_hosts to use the helper for
pooled resources based on utilization
- Enrich active workloads with get_resource_usage() in get_all_active_workloads
- Change resource_type from "memory" to "ram" during host enrollment
- Add GET /workload_hosts/<id>/utilization endpoint with UUID validation
and error handling
Replace hardcoded "worker_agent" string with dynamic WORKER_ID from settings
across Docker monitoring, client operations, and container tasks. This enables
unique identification for multi-worker environments, preventing container
management conflicts.
Implemented a playbook-driven external smoke test framework and initial scenarios with DNS-inclusive checks.
What was added/changed
New files:
smoke/config.py – playbook schema, env overrides, loader, scenario filtering
smoke/runner.py – runner with three scenarios, waiters, DNS/TCP/HTTP probes, JSON reporting
smoke/playbook.example.yaml – sample playbook using IP-based defaults with DNS checks
smoke/README.md – usage and operation docs
smoke/init.py – package marker
API client uplift:
Added IaaSClient.container_lifecycle() to call POST /workloads/containers/<id>/lifecycle/<action> for future restart tests
Leveraged existing client methods:
IaaSClient.get_container() and IaaSClient.delete_container()
IaaSClient.get_container_pods() and IaaSClient.get_container_pod()
Server routes/flow relied on:
Create (async, DB-first): /workloads/containers
Worker status update: update_container_workload()
Pod introspection with DNS and port_forwardings: get_pods(), get_pod()
Allocation + DNS/Tunnel creation: allocate_and_dispatch(), allocate_ports_for_container(), CloudflareTunnelManager
Scenarios implemented (playbook-driven)
create-nginx
Creates a container in a new pod (nginx:stable-alpine by default), requests port 80 with use_dns: true.
Waits for allocation then running.
Discovers ingress from pod.port_forwardings (dns_record_hostname or external_ip_address:external_port).
Probes DNS resolution, TCP connectivity, and HTTP GET / expecting 200 and “nginx”.
Teardown deletes the container, then validates:
TCP connectivity to previously known ingress fails
Port forwarding entries for the container are removed from the pod
Code: _scenario_create_container()
add-second-container
Adds a container to the pod created above, waits running, and probes the new container.
Also probes original container still serves traffic.
Deletes only the newly added container.
Code: _scenario_add_to_pod()
negative-image
Uses a bogus image (thisdoesnotexist.invalid:never), expects final status launch_failed.
Validates no ingress is assigned.
Cleans up the failed container record.
Code: _scenario_negative_image()
Runner capabilities
DNS checks are enforced based on playbook defaults; uses DNS if available and also probes external IP:port as fallback unless dns_required is strictly enforced.
Robust waiters with polling and timeouts: allocation, running, deletion.
Data-plane probes:
DNS resolution: dns_resolve()
TCP connect: tcp_connect()
HTTP GET with redirects and content assertions: http_get()
JSON report with timings and per-endpoint assertions written to configurable path.
How to run
Set env vars or edit the playbook:
API_BASE_URL (default http://172.17.0.1:5000/api)
API_KEY
VDC_ID (target UUID)
DNS_REQUIRED=true (enforced in example playbook)
Optional: WEBSOCKET_SERVER_URL for documentation purposes
Review and update smoke/playbook.example.yaml.
Execute:
python smoke/runner.py --playbook smoke/playbook.example.yaml --report smoke/out/report.json
The runner prints summary and writes detailed JSON results.
Notes and extension points
Container restart validations can leverage IaaSClient.container_lifecycle() and assert data-plane blip/resume while worker transitions via container_lifecycle_action().
Additional coverage can be added for volumes, networks, VMs, and pod lifecycle actions reusing the same framework.
The framework stays external: only public API calls and real-world network probes; no internal DB calls.
This delivers the initial smoke test infrastructure and the required test coverage: create/delete container with DNS and data-plane verification, add container to an existing pod with verification for both, and a negative “dodgy image” test that must fail correctly.
What changed
Unified two-phase workflow across all scenarios (new pod, add to existing pod, and waiting-host retry):
Phase A (Persist): persist_pod_and_containers()
Phase B (Allocate+Dispatch): allocate_and_dispatch()
Celery task updated to always use the unified flow:
process_workload() now calls persist_pod_and_containers() then allocate_and_dispatch() for container requests.
Deprecated legacy flows replaced with thin forwarding wrappers (kept temporarily for compatibility):
_create_new_pod_flow() forwards to persist+allocate.
_add_to_existing_pod_flow() forwards to persist+allocate.
_build_pod_objects_and_enqueue() forwards to persist+allocate.
allocate_and_dispatch refactored to be safe and idempotent for mixed pods:
Only processes pending-allocation containers via _pending_allocation_containers(), preventing status transition errors and duplicate PF/DNS when appending to a running pod.
Extracted small helpers to keep logic simple and testable:
Host selection: select_host_for_pod()
Waiting-host backoff/jitter: set_waiting_host_and_next_retry()
NS controller alignment and launch params: ensure_nscontroller_on_host()
Port allocation and PF creation: allocate_ports_for_container()
Cloudflare tunnel and DNS: ensure_tunnel_and_dns()
External port selection (retained): find_available_external_port()
Status bulk update (retained): _set_container_statuses()
Cloudflare config standardized:
All references now use CLOUDFLARE_API_TOKEN, CLOUDFLARE_ACCOUNT_ID, CLOUDFLARE_ZONE_ID in allocate_and_dispatch() and ensure_tunnel_and_dns().
Why this satisfies the goal
Single workflow for all container creations and management:
New pod creation and adding to an existing pod both start by persisting to DB and then run allocation+dispatch through the same code path allocate_and_dispatch().
Celery retry for waiting-host continues calling allocate_and_dispatch() which now encapsulates selection, PF/DNS, and dispatch.
Safe status transitions and idempotency:
Only new/pending containers are set to allocated on success; existing running containers are untouched, preventing validation errors in Workload._validate_status_change().
Port mappings and DNS are only created for the new/pending containers to avoid duplicates.
Notes and next steps (tracked in TODOs)
Optional further refinement: extract dispatch and status-setting blocks into helpers for even tighter separation.
Add tests covering new pod, add-to-existing, waiting-host retry, and idempotent re-dispatch behavior.
Update docs and curl samples to show both payload shapes.
Key entry points to review
Persist phase: persist_pod_and_containers()
Allocate+Dispatch: allocate_and_dispatch()
Pod payload builder: build_pod_payload()
Celery task: process_workload()
Retry beat: retry_waiting_host_pods()
This achieves a consolidated, maintainable workflow without introducing a single large function, using small helpers to keep concerns separated and the codebase easier to evol
What changed:
Added a context-based ID source using Python ContextVar, used when explicitly set.
Preserved Flask behavior (reads g.request_id when a request context exists).
Added Celery awareness: when running in a Celery worker/task, it uses the current task’s id (or correlation_id) automatically.
Final fallback now includes process/thread info instead of “no-context”.
Resolution of the reported issue:
Your Celery worker logs will now include RequestID set to the Celery task id (or correlation_id). If not in a Celery task and not in Flask, the logger will emit a structured fallback like pid:1234|thr:MainThread rather than no-context.
For non-Flask, non-Celery jobs:
At the beginning of your job or script, set a correlation/request ID via the logger module’s context setter and clear it when done. This ensures all logs in that execution path include your chosen ID.
No changes required to your existing Flask code paths, and the formatter/handlers remain intact, so log output format is unchanged aside from improved request ID values
This commit implements the following changes:
- Add SHAREDFS_ROOT configuration option with default value "/mnt/shared"
- Create _process_storage_volumes method to handle NFS volume directory creation
- Update container launch and validation to support NFS volume mounting
- Ensure storage volumes are processed before container creation in pod update
This commit introduces new API endpoints for retrieving audit logs:
- `/api/audit/container/<container_id>` to get container-specific logs
- `/api/audit/pod/<pod_id>` to get pod-specific logs
- `/api/audit/actions` to list available audit action types
The implementation includes:
- Comprehensive error handling
- Pagination support
- Filtering by action type
- Standard API response formatting
- Detailed logging for debugging
The only thing this doesnt handle is when a client goes into offline due to a transient link failure but it comes back online. We want to avoid updating the DB every time a pong is RX'd.. Perhaps we do a seocndary corss checks with offline hosts?