The change switches from using environment variables to Flask's application configuration for retrieving the WebSocket server URL, ensuring consistent configuration access through the application context.
Dump container console logs when NSController failures or health check failures occur to help with troubleshooting. The new method fetches the last 100 lines of logs and logs them at debug level for easier diagnosis of container startup issues.
Called in two scenarios:
- When NSController failure is detected
- When NSController health check fails
Track NSController and regular container launch failures when health checks fail or containers are skipped. Only send delta status updates when there are no launch failures, preventing misleading status communications.
When NSController health check fails, properly clean up the container instead of just aborting. This prevents resource leaks by stopping and removing the failed container, cleaning up network namespace symlinks, and removing injected files.
Move NetworkPort creation earlier in the allocation process and use the NSController's attached NetworkPort IP address for port forwarding when available, instead of defaulting to the host's IP address. This ensures pods with NSControllers use their dedicated network port IP for external connectivity.
Collects network ports from all containers within a pod and includes
them in the API response. Each network port entry contains full details
including ID, name, network ID, IP address, MAC address, port type,
DNS servers, subnet mask, and status information.
NetworkPort objects are now attached to the NSController (not individual containers).
The NSController manages the network namespace for the entire pod, and containers
within that pod share this network configuration. This design allows containers
to share the same network stack while maintaining proper isolation via the
NSController's network namespace.
The refactor consolidates network port creation by:
- Scanning all containers in a pod to collect network requirements
- Creating NetworkPort records attached to the NSController
- Updating port_type from "container" to "nscontroller"
- Removing explicit networks configuration from NSController config (relying on inherited pod networking)
Extract duplicate deletion code from pod and container routes into
reusable helper functions in deletion_helpers.py module:
- collect_and_cleanup_network_ports() for network port cleanup
- cleanup_workload_related_records() for resource/volume cleanup
- log_deletion_audit_event() for consistent audit logging
Add SDN network port updates during pod deletion to ensure proper
network cleanup when containers are removed.
Optimize get_pods() endpoint by replacing N+1 queries with batch
queries and lookup maps for nscontrollers, tunnels, and DNS records.
Enhance get_pod() response to include cloudflare tunnel configuration
and DNS records for better visibility into pod networking setup.
- Add network port configuration during container creation using OVS_SDN
- Clean up OVS bridge ports when deleting containers
- Import OVS_SDN module for port management operations
- Update delete_container method to accept and process network_ports parameter
Add network port information to pod payload construction by querying
NetworkPort records for each container and including network details
(vni, ovs_bridge, gateway, ip_address, mac_address, dns_servers) in the
container specification.
Add a new delete_port method to the OVS_SDN worker class that enables removing network ports from Open vSwitch bridges. The method includes safety checks to verify port existence before deletion and provides proper error handling and logging for troubleshooting.
Implement comprehensive network port management across container lifecycle operations:
- Add NetworkPort creation when containers are provisioned
- Send SDN updates after port creation for network configuration
- Properly clean up network ports and send SDN updates on container deletion
- Filter out deleted ports from SDN payload calculations and network port queries
- Add dependency on SDN update tasks before container dispatch
These changes ensure proper network state synchronization between the database and SDN controller, preventing stale port configurations and ensuring consistent network topology across container lifecycle operations.
Added validation to skip peer ports with missing workload_host_ip during SDN
payload generation. When a port's peer host has no east-west IP configured, the
system now logs a warning and continues processing rather than allowing the None
value to cause VXLAN interface failures.
Adds a debug endpoint to preview SDN payloads without sending to workers, enabling easier troubleshooting of network configurations. Also fixes indentation issue in generate_sdn_payload that prevented proper return of payload data, and removes blocking debug code that was preventing SDN updates from being processed.
- Add address=/localhost/127.0.0.1 to DNS zone generation for Docker healthchecks
- Configure container DNS to use 127.0.0.1 for NSControllers with empty dns_opt and search
- Update NSController image to xcloudify-nscontroller:latest
- Adjust Dockerfile healthcheck start-period from 10s to 5s for faster startup detection
Adds a comprehensive bash script that demonstrates DNS functionality for containers:
- Creating containers with DNS-enabled ports (automatic DNS record creation)
- Creating custom A and CNAME records
- Listing, updating TTL, and deleting DNS records for a VDC
Also improves pod naming uniqueness in launch_single_container.sh by adding timestamps, and adds nscontroller image to the build script.
Correct the SRV record format in dnsmasq configuration from
<service>,<priority>,<weight>,<port>,<target> to
<service>,<target>,<port>,<priority>,<weight>. Also add support for pods
running in hostNetwork mode by using the host's ip_address_northsouth
when a pod has no network port assigned.
When an alternate host doesn't exist for a pod that's failing on an offline
host, we need to clear the workload_host_id reference so that allocate_and_dispatch
will select a new host instead of repeatedly trying the same offline host.
Refactor NSController configuration into a single source of truth module, ensuring consistent DNS volume mounts and resource limits. Fixes bug where trailing periods in universe DNS names caused double periods in VDC zone names. Removes legacy DNS update modes and standardizes on shared volume approach. Adds comprehensive configuration constants and improves worker-side DNS handling.
Add ENABLE_WEBSOCKET_PING_DEBUG setting to control whether websocket ping
diagnostic logs are emitted. This reduces log noise during normal operation
while allowing detailed ping debugging when needed.
Enhance the get_container_workload endpoint to include resource usage
data in the response. Additionally adds validation for the optional
pod_name parameter to ensure non-empty string values.
Remove the periodic_connection_monitor method and its associated task
lifecycle management. This eliminates the 10-second interval health
checks that logged transport state and detected split-brain, stale
flag, and ghost state conditions.
- Add periodic health monitor to detect split-brain and ghost connection states
- Disable automatic reconnection in favor of manual control
- Fix reconnection_in_progress flag not being reset after successful reconnect
- Add structured logging with prefixes ([RECONNECT], [HEALTH], [PING], [START])
- Log transport state during connection lifecycle events for better debugging
Implement PCI device scanner to detect and report GPUs, TPUs, and other
hardware accelerators on workload hosts at startup. The scanner follows
the same pattern as the existing OVS bridge scanner.
Key components:
- PCIDeviceScannerTask: Scans devices using lspci, filters based on
vendor/device ID combinations from region configuration, and reports
to API server
- Worker integration: Scanner initializes on worker startup and executes
on join accept
- Region-based filtering: JSON configuration allows per-region device
filters with wildcard support
- Storage: Uses existing WorkloadHostFixedResource model with type
'gpu', 'tpu', or 'accelerator'
The scanner loads filters from region config via API, executes lspci,
parses device information, applies filters, and reports matched devices
to the API endpoint /api/workload_hosts/{host_id}/pci_devices.
Includes comprehensive documentation covering implementation details,
configuration examples, testing procedures, and troubleshooting guidance.
Add Ansible playbook for deploying nginx webserver and content generator
using shared NFS storage. Update build script to include new container image,
add port mapping to launch example, and note related todos for reallocation
and NSController CPU issues.
Add comprehensive support for pod management (create/attach/lifecycle), healthchecks (test/interval/timeout/retries/start_period), restart policies (always/on-failure/etc. with max retries), injected files (base64 content/permissions), and partial updates (PATCH for zero-downtime env/resources changes). Update SDK with new methods (patch_container, pod_lifecycle, wait_for_* with backoff), enhanced auth (token refresh), validation, and error handling. Include updated examples, tests, docs, and schemas for backward compatibility and 100% feature coverage.
Bump role version to 2.0.0; deprecate legacy full-recreate with warnings.
- Replace multiprocessing.Process with threading.Thread for Docker and Libvirt monitors
- Remove multiprocessing.Manager usage and associated cleanup code
- Simplify monitor lifecycle with running flag instead of Event objects
- Add use_db_state=True consistently to all build_pod_payload calls
- Add pod_id field to containers in workload host routes
- Add debug logging for host status updates and container data
Threading approach reduces complexity and overhead while maintaining
the same monitoring functionality.
Adds logic to set object_type and object_id when object is None,
preventing AttributeError in batch operations. Updates debug logging
and AuditEntry creation to use these variables.
- Add restart_policy field to ScenarioSpec class in config.py
- Update runner.py to conditionally include restart_policy in container specs for create_container, add_container_to_pod, and create_container_with_invalid_image methods
- Add example scenario with restart_policy in playbook.example.yaml
- Create test_restart_policy.yaml playbook for testing restart policy functionality
- Create test_restart_policy_direct.py script for direct API testing of restart policy feature
- Add validation for 'restart_policy' parameter in workload container routes with valid options: 'no', 'always', 'unless-stopped', 'on-failure'
- Set default restart policy to 'always' if not specified
- Update pod payload builder to include restart policy from launch params, defaulting to 'always'
- Modify container task to use restart policy from container spec or settings, with logging for applied policy
Implements configurable container restart behavior as per recent config addition.
Adds two new Ansible playbooks demonstrating deployment of Node-RED and TP-Link Omada Controller to xCloudify. Each example includes container configuration, volume definitions, and deployment summaries.
- node-red.yml: Deploys Node-RED with configurable module installation and authentication
- omada-controller.yml: Deploys Omada Controller with required ports and storage volumes
Both examples use the xcloudify_infrastructure role and include debug output for deployment verification.
Add "docker_absent" to the list of handled Docker events and map it to the "deleted" status in the event-to-status dictionary. This ensures proper handling of container absence events.
- Introduce update_pod_status function to synchronize pod status with container statuses
- Set pod to 'mixed' if containers have varying statuses, or match uniform status
- Soft delete pod if all containers are deleted
Closes#123 (assuming related issue)
Add 'include_deleted' query parameter to the get_container_workloads_for_host endpoint,
allowing clients to retrieve deleted containers in the response. Update build_pod_payload
function to accept and handle the include_deleted flag, filtering workloads accordingly.
This enhances the API's flexibility for reconciliation tasks by providing access to
deleted container data when needed.
Add validation for injected_files parameter in API, including security checks for filenames, base64 content, and permissions. Update pod payload building to include injected files. Implement file injection logic in worker tasks, with base64 decoding, temp file creation, bind mounts, and cleanup. Include comprehensive test script for validation and handling.
Make report_container_statuses and fetch_container_workloads_for_host properly async,
ensuring correct handling of asynchronous operations in the worker client.
The get_container_workloads endpoint now supports an optional 'include_deleted' query parameter. When set to 'true', it includes deleted container workloads in the response. Defaults to 'false' for backward compatibility.
Add logic to fetch current desired container workloads from the server and filter status reporting to only include containers that should exist according to the server response. This prevents reporting on stale or unexpected containers, improving accuracy and reducing noise in status updates.
- Add new single-container.yml example for deploying nginx container
- Filter out deleted/pending-deleted containers in management operations
- Enhance error handling in network module with safer result access
- Validate provided VNI for uniqueness within the region
- Generate random VNI if not provided
- Improve error handling for VDC not found and VNI conflicts
BREAKING CHANGE: VNI is now required to be unique per region, may affect existing networks with duplicate VNIs
- Extract original pod update logic into _process_pod_update method
- Introduce handle_pod_update_with_reconciliation for blacklist-based noise reduction and status reconciliation
- Add methods to capture, compare, and send delta container status updates
- Implement blacklisting/unblacklisting to prevent event processing during updates
- Update handle_pod_update to delegate to new reconciliation approach
- Remove NSController recreation logic from core update flow
- Add shutdown method for graceful Docker connection closure
This enhances pod update reliability by reducing event noise and ensuring accurate status reporting through reconciliation.