autostart mds in nscontroller
added missing docker volume/packages for socket mounts for ns
added network id env for nscontroller
add missing module of sdn_helper
error handle when ovs-ofctl is empty
Added APIs to expose the metadata endpoints (from api server)
Proxy in worker hosts that listen on sockets and forwards metadata requests to API
Proxy in NSController to forward metadata requests via socket to its worker host.
API to register GPU (worker does at initialization), list all GPUs (to fetch during creation menu)
Schedular to Match with relevant hosts filter
A Table to store GPU models
A table to map gpu to pcie to workload and host
Dynamic binding unbinding of vfio, nvidia driver at runtime for passthrough
Implement NS Controller API for Dispatch & Delete to Per Host Per Network
Chore: Used_ips filter to show deleted (network.py:106) while creating new
DHCP templated w DNSMasq (Mapped Port to IP in NSController)
Protect NSControllers from Delete during Reconsilation
Warning! : It uses Static, Hardcoded at libvirt_xml_builder.py
Implemented GPUCapableHosts Filter
Santize GPU Payload from User Input
Add Endpoint to register PCIE Devices
Seed the all regions with pcie filter of AMD and NVIDIA GPU
Previous code accessed resource.resource_type and resource.quantity which
never existed on WorkloadHostFixedResource. The correct fields are type,
model, manufacturer and physical_address, matching the model definition
and DB schema since the initial migration.
Bug introduced in 1731e2c.
VM:
Enabled Filter VM Scheduling by HasSufficientResources
Implemented to Record Resource Usage by WorkloadResourceUsage; creates new record on basis of resource_type (CPU,RAM)
Volume:
Implemented Local Volume vm spin (Fallback Local and reject other volume types for spinning vm),
Fix: hardcode to image for image backed volume in vm payload :298
Network:
Added error handle for port creation, bridge heirarchy selection (network to stored)
The change switches from using environment variables to Flask's application configuration for retrieving the WebSocket server URL, ensuring consistent configuration access through the application context.
Add XCF_DEFAULT_OVS_BRIDGE environment variable to configure a default
OVS bridge for networks without explicit configuration. Update payload
builders to fallback to this default when generating SDN and pod
configurations.
- Add default_ovs_bridge column to regions table via migration
- Update network API routes to support ovs_bridge field
- Implement fallback logic in WorkloadHost and pod payload builders
- Adjust NSController resource allocation (4mCPU/128MiB) and image
Add VDC-specific DNS configuration directories to prevent cross-tenant
DNS leakage. Each VDC now gets its own subdirectory under the global DNS
config path.
Introduce sanitize_dns_name() function to handle edge cases in DNS names
including empty strings, trailing dots, .local suffix duplication, and
consecutive periods.
Add environment variable WEBSOCKET_SERVER_URL for WebSocket server
configuration with fallback to default.
Update dnsmasq configuration to enable expand-hosts and improve DNS
rebind protection.
Add validation to ensure network IDs referenced in container configurations
actually exist and belong to the specified VDC. This prevents processing
containers with invalid or cross-VDC network references, improving data
integrity and providing clear error messages for misconfigured workloads.
Move NetworkPort creation earlier in the allocation process and use the NSController's attached NetworkPort IP address for port forwarding when available, instead of defaulting to the host's IP address. This ensures pods with NSControllers use their dedicated network port IP for external connectivity.
Collects network ports from all containers within a pod and includes
them in the API response. Each network port entry contains full details
including ID, name, network ID, IP address, MAC address, port type,
DNS servers, subnet mask, and status information.
NetworkPort objects are now attached to the NSController (not individual containers).
The NSController manages the network namespace for the entire pod, and containers
within that pod share this network configuration. This design allows containers
to share the same network stack while maintaining proper isolation via the
NSController's network namespace.
The refactor consolidates network port creation by:
- Scanning all containers in a pod to collect network requirements
- Creating NetworkPort records attached to the NSController
- Updating port_type from "container" to "nscontroller"
- Removing explicit networks configuration from NSController config (relying on inherited pod networking)
Extract duplicate deletion code from pod and container routes into
reusable helper functions in deletion_helpers.py module:
- collect_and_cleanup_network_ports() for network port cleanup
- cleanup_workload_related_records() for resource/volume cleanup
- log_deletion_audit_event() for consistent audit logging
Add SDN network port updates during pod deletion to ensure proper
network cleanup when containers are removed.
Optimize get_pods() endpoint by replacing N+1 queries with batch
queries and lookup maps for nscontrollers, tunnels, and DNS records.
Enhance get_pod() response to include cloudflare tunnel configuration
and DNS records for better visibility into pod networking setup.
Add network port information to pod payload construction by querying
NetworkPort records for each container and including network details
(vni, ovs_bridge, gateway, ip_address, mac_address, dns_servers) in the
container specification.
Implement comprehensive network port management across container lifecycle operations:
- Add NetworkPort creation when containers are provisioned
- Send SDN updates after port creation for network configuration
- Properly clean up network ports and send SDN updates on container deletion
- Filter out deleted ports from SDN payload calculations and network port queries
- Add dependency on SDN update tasks before container dispatch
These changes ensure proper network state synchronization between the database and SDN controller, preventing stale port configurations and ensuring consistent network topology across container lifecycle operations.
Added validation to skip peer ports with missing workload_host_ip during SDN
payload generation. When a port's peer host has no east-west IP configured, the
system now logs a warning and continues processing rather than allowing the None
value to cause VXLAN interface failures.
Adds a debug endpoint to preview SDN payloads without sending to workers, enabling easier troubleshooting of network configurations. Also fixes indentation issue in generate_sdn_payload that prevented proper return of payload data, and removes blocking debug code that was preventing SDN updates from being processed.
- Add address=/localhost/127.0.0.1 to DNS zone generation for Docker healthchecks
- Configure container DNS to use 127.0.0.1 for NSControllers with empty dns_opt and search
- Update NSController image to xcloudify-nscontroller:latest
- Adjust Dockerfile healthcheck start-period from 10s to 5s for faster startup detection
Correct the SRV record format in dnsmasq configuration from
<service>,<priority>,<weight>,<port>,<target> to
<service>,<target>,<port>,<priority>,<weight>. Also add support for pods
running in hostNetwork mode by using the host's ip_address_northsouth
when a pod has no network port assigned.
When an alternate host doesn't exist for a pod that's failing on an offline
host, we need to clear the workload_host_id reference so that allocate_and_dispatch
will select a new host instead of repeatedly trying the same offline host.
Refactor NSController configuration into a single source of truth module, ensuring consistent DNS volume mounts and resource limits. Fixes bug where trailing periods in universe DNS names caused double periods in VDC zone names. Removes legacy DNS update modes and standardizes on shared volume approach. Adds comprehensive configuration constants and improves worker-side DNS handling.
When a worker is marked offline during reconciliation, automatically
schedule the host downtime handler to ensure proper workload
rescheduling. The handler is triggered with a configurable countdown
delay from HOST_DOWNTIME_CONFIRM setting.
This ensures that workloads are properly handled when workers become
unavailable, improving system reliability and task management.
Add external and internal IP addresses, DNS record details, and filter out deleted port forwardings. Include null checks for workload and DNS record.
BREAKING CHANGE: Replaced "ip_address" with "external_ip_address" and added "internal_ip_address" in port forwarding response.