Commit Graph
100 Commits
Author SHA1 Message Date
coryHawkvelt a71c74f133 Merge pull request 'Implement Openstack Style Metadata with Network ID population' (#49) from JamesBhattarai/theapi:v1.01/network into master
Reviewed-on: coryHawkvelt/theapi#49
2026-05-25 03:35:37 +00:00
coryHawkvelt 28170a2650 Merge pull request 'Feat: Implemented LVM API & Debug Volume Fits Image' (#43) from JamesBhattarai/theapi:v1.01/volumes-xcloudify into master
Reviewed-on: coryHawkvelt/theapi#43
2026-05-13 04:40:29 +00:00
coryHawkvelt 7778e5151f Merge pull request 'Feat: Implemented Logging Stream for VM's Log (Via Libvirt Console)' (#42) from JamesBhattarai/theapi:v1.01/logs-vm-xcloudify into master
Reviewed-on: coryHawkvelt/theapi#42
2026-05-13 04:40:19 +00:00
coryHawkvelt 433bec4556 feat(xcloudify): add ovs_bridge support for network management
Add ovs_bridge parameter to xcloudify_network module and client utilities
to enable OVS bridge configuration for VXLAN networks. Includes validation,
documentation, and example updates.
2026-01-27 10:44:02 +10:30
coryHawkvelt a6eb0eb4b7 feat(network): add ovs bridge fallback configuration
Add XCF_DEFAULT_OVS_BRIDGE environment variable to configure a default
OVS bridge for networks without explicit configuration. Update payload
builders to fallback to this default when generating SDN and pod
configurations.

- Add default_ovs_bridge column to regions table via migration
- Update network API routes to support ovs_bridge field
- Implement fallback logic in WorkloadHost and pod payload builders
- Adjust NSController resource allocation (4mCPU/128MiB) and image
2026-01-27 10:42:02 +10:30
coryHawkvelt bdceb98b5f feat(xcloudify):ansible-example- resolve network names to ids
Add network name resolution functionality to allow users to specify
network names instead of IDs in container configurations. The client
now resolves network names to their corresponding IDs using the
networks dictionary, falling back to using the value as-is if the
name is not found (supporting direct ID usage).
2026-01-25 06:21:54 +10:30
coryHawkvelt 62243c4e4b feat(dns): add vdc dns isolation and improve sanitization
Add VDC-specific DNS configuration directories to prevent cross-tenant
DNS leakage. Each VDC now gets its own subdirectory under the global DNS
config path.

Introduce sanitize_dns_name() function to handle edge cases in DNS names
including empty strings, trailing dots, .local suffix duplication, and
consecutive periods.

Add environment variable WEBSOCKET_SERVER_URL for WebSocket server
configuration with fallback to default.

Update dnsmasq configuration to enable expand-hosts and improve DNS
rebind protection.
2026-01-25 06:21:20 +10:30
coryHawkvelt 722b22a5ab chore(api): remove debug logging from get_pods endpoint 2026-01-25 05:39:02 +10:30
coryHawkvelt 69aafb39ca feat(nscontroller): wait for health check before launching pods
Implements health check synchronization to ensure NSController is fully
operational before launching dependent containers in a pod.

- Add wait_for_nscontroller_health() method to poll container health status
- Create dedicated healthcheck.sh script for faster process validation
- Optimize health check intervals (5s) and timeout (3s) for quicker detection
- Update dnsmasq.conf with DNS rebind protection
- Abort pod launch if NSController health check fails within timeout (120s)
2026-01-25 04:37:20 +10:30
coryHawkvelt 78d06d5647 Add iperf and tcpdump to nscontroller 2026-01-24 16:45:15 +10:30
coryHawkvelt e8901ded98 feat(worker): improve network namespace handling with PID-based symlinks
Refactor network namespace management to use container PID directly instead of Docker's SandboxKey. Changes include:

- Use /proc/<pid>/ns/net path for creating namespace symlinks in /run/netns/
- Create /run/netns directory if it doesn't exist for better robustness
- Add /proc:/proc:ro volume mount to docker-compose for namespace access
- Support configurable OVS bridge in delete_port operations (default: br-int)
- Clean up symlinks directly instead of using ip netns commands

This approach provides more reliable namespace access and better aligns with standard Linux namespace tooling.
2026-01-24 15:31:05 +10:30
coryHawkvelt ceecc15ac7 feat(worker): add network namespace directory mounts
Mount /run/netns and /var/run/docker/netns to give the worker
container access to host and Docker network namespaces, enabling
namespace management operations.
2026-01-21 21:27:31 +10:30
coryHawkvelt 7077644ed2 feat(worker): add configurable bridge support to ovs worker tasks
Remove hardcoded "br-int" bridge references and implement per-network
configurable bridge support across OVS SDN and libvirt components.

- Replace BRIDGE constant with required ovs_bridge parameter
- Add bridge parameter to all OVS operations (flows, meters, ports, vxlan)
- Enable per-bridge flow processing and cleanup
- Add validation raising ValueError when ovs_bridge parameter missing
- Track bridges in use and process each independently
- Update libvirt_xml_builder to extract and validate ovs_bridge from config
- Add comprehensive docstrings and improve logging

BREAKING CHANGE: ovs_bridge parameter is now required in all network configurations. Configurations missing this parameter will raise ValueError.
2026-01-21 21:23:14 +10:30
coryHawkvelt b00c2a7e00 feat(models): add ovs_bridge field to port configuration output 2026-01-21 18:44:32 +10:30
coryHawkvelt 22beef3f64 fix(worker): make sudo usage conditional for network namespace commands
Add _should_use_sudo() helper function to detect when running inside Docker
where sudo is not available. Update all network namespace operations to use
conditional sudo prefix instead of hardcoded sudo commands. This ensures the
code works correctly in both Docker and non-Docker environments.
2026-01-21 14:33:17 +10:30
coryHawkvelt 857e2604de fix(worker): enhance NSController network mode validation to support bridge mode
NSControllers now use bridge network mode when no OVS bridges are configured, allowing Docker port publishing to work correctly. The validation logic checks for OVS bridge presence in network_ports and expects "none" mode only when OVS ports will be added, otherwise expects "bridge" mode to match launch_container behavior.
2026-01-21 07:46:40 +10:30
coryHawkvelt e13f18d14c feat(worker): allow NSControllers to use Docker bridge when OVS not configured
NSControllers previously always used network_mode=none for OVS port attachment.
This change adds logic to check if OVS bridge is configured, and only uses
network_mode=none when OVS ports are present. Otherwise, NSControllers can use
Docker bridge for standard port publishing.
2026-01-21 06:44:53 +10:30
coryHawkvelt ffa1f45f08 feat(worker): add OpenVSwitch support to worker containers
Add openvswitch-switch package to Docker images and mount OVS socket directories to enable OVS port attachment functionality for container namespaces.
2026-01-21 05:18:09 +10:30
coryHawkvelt 8482018145 feat(worker): add OVS port attachment to container namespaces for NSControllers
This feature enables NSControllers with network_mode="none" to have OVS ports attached directly to their network namespace. The implementation includes:

- attach_ovs_ports_to_container_namespace: Creates namespace symlinks, moves OVS ports into the container's network namespace, and configures them with the specified MAC and IP addresses

- cleanup_container_namespace: Cleans up network namespace symlinks when containers are deleted

This is needed for NSControllers that require direct OVS port integration while using network_mode="none" for isolation.
2026-01-21 04:50:35 +10:30
coryHawkvelt 3bb4d92903 fix(worker): set NSController network mode to none
NSControllers should use network_mode="none" instead of Docker bridge networking, as they receive OVS ports directly. This fix ensures NSControllers skip Docker network attachment and validates the network mode configuration.
2026-01-21 04:31:37 +10:30
coryHawkvelt 51f68560a4 fix(worker): skip status sync when launch failures exist
When launch failures occur, the launch_failed event is already sent separately to handle the status update. Sending a full status sync in this case is redundant and could cause confusion. This change adds a check to skip the full status sync when there are launch failures, only logging the number of failures instead.
2026-01-21 04:16:21 +10:30
coryHawkvelt 364e7177fe fix(api): validate network references in container workloads
Add validation to ensure network IDs referenced in container configurations
actually exist and belong to the specified VDC. This prevents processing
containers with invalid or cross-VDC network references, improving data
integrity and providing clear error messages for misconfigured workloads.
2026-01-17 15:57:03 +10:30
coryHawkvelt d57bfec6c5 fix(worker): improve network port configuration error handling and cleanup
Add proper error handling for OVS network port configuration failures with
container cleanup on error. Changed OVS port creation to raise exceptions
instead of silently continuing, and added bridge parameter validation. When
port configuration fails, the container is now stopped and removed to prevent
orphaned resources.
2026-01-17 15:39:48 +10:30
coryHawkvelt 17f486a694 refactor(container): improve container state logging and reporting
Enhanced debug logging to show desired_state and specific actions being taken
for both NSController and regular containers. Changed event reporting from
'die' to 'absent' for missing containers to more accurately reflect their state.
2026-01-17 15:26:14 +10:30
coryHawkvelt 9454e70f16 Merge pull request 'Anthropics take on the DNS challenge' (#28) from DNS-anthropic into master
Reviewed-on: coryHawkvelt/theapi#28
2025-12-09 03:46:25 +00:00
coryHawkvelt 0486ce97b8 update todo 2025-12-04 17:15:34 +10:30
coryHawkvelt d0143627fc Anthropics take on the DNS challenge 2025-12-04 10:05:10 +10:30
coryHawkvelt c9f7be8ab9 revised nscontroller to include DNSmasq 2025-12-04 09:43:32 +10:30
coryHawkvelt 0a3371262e DNS Docs v2 2025-12-04 09:12:37 +10:30
coryHawkvelt be89b88e67 DNS Proposal 2025-12-04 07:34:18 +10:30
coryHawkvelt 00e04d69be feat(tasks): schedule downtime handler when marking workers offline
When a worker is marked offline during reconciliation, automatically
schedule the host downtime handler to ensure proper workload
rescheduling. The handler is triggered with a configurable countdown
delay from HOST_DOWNTIME_CONFIRM setting.

This ensures that workloads are properly handled when workers become
unavailable, improving system reliability and task management.
2025-12-04 00:43:46 +10:30
coryHawkvelt b989fa9068 feat(infra): implement pod update logic and fix container matching
- Add _pod_needs_update method to determine when pods require recreation
- Improve container name matching to support both 'name' and 'container_name' keys
- Update xcloudify_container module to delete and recreate pods when updates are needed
- Fix missing Jinja2 loop close in simple-web-app example
- Disable use_dns in example configuration
2025-12-04 00:22:16 +10:30
coryHawkvelt 691f845529 feat(api): enhance pod response with detailed port forwarding info
Add external and internal IP addresses, DNS record details, and filter out deleted port forwardings. Include null checks for workload and DNS record.

BREAKING CHANGE: Replaced "ip_address" with "external_ip_address" and added "internal_ip_address" in port forwarding response.
2025-12-03 23:29:50 +10:30
coryHawkvelt adf23c4a84 feat(infra): add support for custom pod names in container deployments
Pod names enable grouping containers into pods with shared networking and lifecycle management. This feature is implemented across Ansible examples, roles, and API scripts, allowing users to specify or create pods by name for better organization of multi-container deployments.
2025-12-03 22:03:38 +10:30
coryHawkvelt c42c352a78 NOT REVIEWED - Failover logic enhancements 2025-11-06 13:28:12 +10:30
coryHawkvelt 042fdb8478 docs and notes 2025-11-06 13:25:40 +10:30
coryHawkvelt 774e5c190b refactor(worker): defer manager initialization to worker processes
move multiprocessing manager initialization from constructor to run() method in both DockerMonitor and LibvirtMonitor to avoid issues with forking managers across processes

fix(worker): prevent stop event errors during shutdown

add null checks for stop_event and manager before attempting to use or shutdown them in monitor classes

refactor(worker): simplify process management in main

replace processes list with direct worker_proc handling for cleaner shutdown logic

feat(worker): add error handling for monitor process startup

wrap monitor process creation in try-except blocks to catch and log startup failures
2025-11-04 15:28:45 +10:30
coryHawkvelt 7fd7cd9aaa feat(worker): implement monitor creation inside worker process and auto-restart
- Move DockerMonitor and LibvirtMonitor instantiation into the worker process after successful join
- Add support for disabling monitors via no_docker and no_libvirt flags
- Implement automatic restart of worker process on unexpected termination
- Add reconnection logic to WorkerClient with exponential backoff
- Ensure container reconciliation runs on both initial join and reconnections
2025-11-04 15:12:09 +10:30
coryHawkvelt 87adb4938b feat(tasks): add offline host check with backoff in allocation retry
Introduce handling for pods with assigned but offline hosts in the retry task.
Increment allocation attempts, apply exponential backoff with jitter, and
schedule next retry accordingly. Skip immediate reallocation if host is offline
to prevent unnecessary attempts. Configuration values are used for base delay,
backoff factor, jitter, and max delay.
2025-11-04 14:41:59 +10:30
coryHawkvelt 8586ace1d6 refactor(container): optimize port allocation and NSController handling for migration
Refactored the container allocation logic to improve migration support:
- Moved NSController handling directly into main allocation flow
- Added migration-aware port allocation with existing port reuse
- Simplified function signatures by removing ns_lp parameter dependency
- Moved DB commit earlier in the process for better transaction management
- Improved stale port mapping cleanup with proper error handling
- Fixed typos in logging messages
2025-11-04 14:02:56 +10:30
coryHawkvelt 1027f77011 feat(worker): add full status sync, map statuses
- Add send_full_statuses_for_host to emit a full snapshot after
  reconciliation for convergence. Supports "host" or "pod" scope via
  WORKER_FULL_STATUS_SCOPE, filters by managed_by label, and skips
  neutral states.
- Map Docker container statuses to websocket_server event types in
  WorkerClient: "running" -> running, "exited"/"dead"/"stopped" -> die,
  skip others. Align details.status and Action accordingly to avoid
  unsupported events.
- Abort launching regular containers if NSController launch fails and
  improve logging/injection behavior.

These changes improve state convergence and robustness while preventing
invalid event emissions.
2025-11-04 11:58:53 +10:30
coryHawkvelt 77140d1b95 fix(pod): reassign host if mismatched in allocation
- Update host assignment condition to reassign if workload host ID is set but does not match the selected host
- Add debug logging for host selection and pre-populated host cases
- Add TODO for setting pod state to allocated after container allocation
2025-11-03 12:42:55 +10:30
coryHawkvelt bbfab8b8b7 docs and examples 2025-10-27 08:42:52 +10:30
coryHawkvelt 0237024d1c feat(task-assigner): add debug flag for task assignment logging
Introduce `ENABLE_TASK_ASSIGNMENT_DEBUG` configuration to control verbose
logging in task assignment, Redis subscription, and worker management flows.
This reduces log noise by default while allowing detailed debugging when
enabled.
2025-10-27 08:42:45 +10:30
coryHawkvelt 54ac191f65 Merge branches 'master' and 'master' of ssh://source.hawkless.id.au:222/coryHawkvelt/theapi 2025-10-27 08:00:56 +10:30
coryHawkvelt 3cd200477e feat(worker): add OVS bridge scanner task
Integrate OVSBridgeScannerTask into WorkerClient to scan for local OVS
bridges using ovs-vsctl, report them to the API endpoint, and handle
initialization with error logging. Includes retry logic for API calls
and graceful fallback if OVS is unavailable.
2025-10-27 07:59:03 +10:30
coryHawkvelt 3070b0c703 feat(api): add endpoints for retrieving and updating OVS bridges
Add GET endpoint to fetch non-deleted OVS bridge names for a workload host.

Add POST endpoint to update OVS bridges: validate input, soft-delete removed
bridges, add new bridges, update timestamps for unchanged ones, with error
handling and logging.
2025-10-27 07:46:08 +10:30
coryHawkvelt 67df40def2 feat(container): add lifecycle state management and rename cpu_shares to cpu
Introduce support for container states (present, absent, started, stopped,
restarted) via xcloudify_containers_state variable in Ansible role.

Update container specs and module to use 'cpu' instead of 'cpu_shares' for
allocation, with validation and API adjustments.

Update examples and tasks for state handling, debug output, and error conditions.

BREAKING CHANGE: 'cpu_shares' renamed to 'cpu' in container configurations and module options.
2025-10-27 07:26:43 +10:30
coryHawkvelt 3bca326b51 feat(websocket): add debug flag for ping operations
Introduce `ENABLE_WEBSOCKET_PING_DEBUG` environment variable to
conditionally enable verbose logging for websocket ping sends and
pong receptions, reducing log noise in production.
2025-10-27 07:21:16 +10:30
coryHawkvelt 05f9862f16 refactor(container): fetch pod via query using pod_id 2025-10-27 07:02:34 +10:30
coryHawkvelt c9d0a1cef7 chore(container): remove commented-out pod mapping code 2025-10-27 06:52:34 +10:30
coryHawkvelt 309b132e9c feat(scheduling): add generic resource filter with overcommit support
Refactor HasSufficientResources to support arbitrary resource types via
host.get_resource_utilization(), applying per-resource overcommit ratios
(default 1.0). Introduce dynamic aggregation of pod resource requirements
from WorkloadResourceUsage database entries, replacing hardcoded minima.

Integrate the filter early in select_host_for_pod pipeline when CPU/RAM
requirements exceed zero, preserving filter order for efficiency.

Adjust global cloudflared sidecar limits: CPU to 1 (from 3), RAM to 64MiB
(from 132MiB) for reduced overhead.
2025-10-27 05:57:44 +10:30
coryHawkvelt 4a3681c667 feat(cloudflare): add reconciliation worker and beat task
- Introduce CloudflareReconciliationWorker to reconcile Cloudflare
  tunnels/DNS with DB and clean up stale resources
- Add Celery task tasks.cloudflare_reconciliation and schedule it
  every 5 minutes via beat
- Provide standalone script to run the reconciliation manually
- Add optional scheduler helper for periodic runs

Also:
- Update tunnel lookup to not ignore tunnels in down status so
  cleanup can find and delete them

No breaking changes.
2025-10-27 04:16:11 +10:30
coryHawkvelt 2d9ee4882d fix(container-deletion): pass workload object to cloudflare cleanup function
- modify `cleanup_cloudflare_resources` to accept workload object instead of container id
- remove redundant workload lookup inside the function
- update usage in `process_container_deletion` to pass container object
- adjust resource usage assignments in `ensure_tunnel_and_dns` to use sidecar id
- remove obsolete `_cleanup_nscontroller_cloudflare_resources` function
2025-10-27 03:33:04 +10:30
coryHawkvelt ff0a235355 feat(workload): introduce global resource limits for cloudflared sidecar
refactor resource limit definitions by using global variables for cpu and memory, and track resource usage in WorkloadResourceUsage entries for the NSController workload
2025-10-27 03:12:47 +10:30
coryHawkvelt 23bddecd21 refactor(deletion): consolidate deletion logic into unified handler
Move all workload and pod deletion logic into a new unified deletion handler module. This includes marking workloads as pending-deleted, cleaning up associated resources (port forwardings, volume mappings, network ports, Cloudflare DNS records and tunnels), and updating pod statuses after container deletions. The change reduces code duplication across multiple files and ensures consistent cleanup behavior.

The new app/utils/deletion_handler.py module contains functions for:
- mark_workload_pending_deleted: marks workloads and associated resources as pending-deleted
- cleanup_workload_resources: performs actual resource deletion
- cleanup_cloudflare_resources_for_workload: handles Cloudflare-specific cleanup
- mark_pod_pending_deleted: marks all containers in a pod for deletion
- update_pod_status_after_deletion: updates pod status based on container states

Modified existing code to use these new consolidated functions instead of repeating similar logic in multiple places.
2025-10-27 01:21:37 +10:30
coryHawkvelt 0aa048ba02 junk 2025-10-27 00:00:53 +10:30
coryHawkvelt ea1c3a6e82 refactor(error-handling): introduce custom exceptions for cloudflare and port errors
Move NoFreePortsError to dedicated exceptions module and add CloudFlareException for
Cloudflare-specific failures. Update ensure_tunnel_and_dns to raise CloudFlareException
on errors. Integrate set_waiting_host_and_next_retry in allocation_retry and
process_workload_request tasks for consistent failure handling with exponential backoff.
Handle pod ID as string in set_waiting_host_and_next_retry and add logging.
2025-10-26 23:53:20 +10:30
coryHawkvelt eb4243b97c Dont override cpu and meory for sidecards, it's explicitly defined when creating them 2025-10-26 23:49:54 +10:30
coryHawkvelt e31fda6220 feat(models): add resource utilization methods for hosts and workloads
Add get_resource_utilization to WorkloadHost to compute available and used
resources dynamically from pooled resources and active workloads.

Add get_resource_usage to Workload to return list of used resources for the
workload.

Fix minor typo in TODO comment and add SQLAlchemy func import.
2025-10-26 23:11:05 +10:30
coryHawkvelt 0989a66198 refactor(container): rename cpu_shares to cpu and update resource defaults
Renamed the CPU resource parameter from 'cpu_shares' to 'cpu' across API
validation, payload building, resource usage tracking, and worker tasks for
consistency and simplicity. Adjusted default mem_limit from 128MB to 129MB.
Introduced global constants for NSController resources (5 CPU, 199MB) and set
tunnel sidecar limits (3 CPU, 257MB). Removed unused import.

BREAKING CHANGE: Container API payloads now expect 'cpu' key instead of 'cpu_shares
2025-10-26 22:53:14 +10:30
coryHawkvelt a15f43d03b feat(api): add host utilization endpoint and enrich responses
Introduce _enrich_host_data helper to integrate resource utilization data
into host responses, replacing manual resource queries.

- Update get_workload_host and get_workload_hosts to use the helper for
  pooled resources based on utilization
- Enrich active workloads with get_resource_usage() in get_all_active_workloads
- Change resource_type from "memory" to "ram" during host enrollment
- Add GET /workload_hosts/<id>/utilization endpoint with UUID validation
  and error handling
2025-10-26 22:52:22 +10:30
coryHawkvelt a9e25fe6a7 refactor(worker): use configurable WORKER_ID for container management
Replace hardcoded "worker_agent" string with dynamic WORKER_ID from settings
across Docker monitoring, client operations, and container tasks. This enables
unique identification for multi-worker environments, preventing container
management conflicts.
2025-10-26 21:08:32 +10:30
coryHawkvelt f8721a49ec minor tidy 2025-10-26 21:08:21 +10:30
coryHawkvelt 9fa32731be ansible role 2025-09-17 20:27:50 +10:00
coryHawkvelt 9df4c965d3 upd todo 2025-08-27 17:28:29 +09:30
coryHawkvelt f115bf05a5 Dont duplicate containers at lunch time 2025-08-27 17:28:22 +09:30
coryHawkvelt a2596b954a default all containers to 1gb a cpu 2025-08-27 17:27:10 +09:30
coryHawkvelt 6005c61cc7 Remove workload_host_id from WorkloadResourceUsage 2025-08-27 16:45:51 +09:30
coryHawkvelt fac6d55ff0 Smoke test
Implemented a playbook-driven external smoke test framework and initial scenarios with DNS-inclusive checks.

What was added/changed

New files:
smoke/config.py – playbook schema, env overrides, loader, scenario filtering
smoke/runner.py – runner with three scenarios, waiters, DNS/TCP/HTTP probes, JSON reporting
smoke/playbook.example.yaml – sample playbook using IP-based defaults with DNS checks
smoke/README.md – usage and operation docs
smoke/init.py – package marker
API client uplift:
Added IaaSClient.container_lifecycle() to call POST /workloads/containers/<id>/lifecycle/<action> for future restart tests
Leveraged existing client methods:
IaaSClient.get_container() and IaaSClient.delete_container()
IaaSClient.get_container_pods() and IaaSClient.get_container_pod()
Server routes/flow relied on:
Create (async, DB-first): /workloads/containers
Worker status update: update_container_workload()
Pod introspection with DNS and port_forwardings: get_pods(), get_pod()
Allocation + DNS/Tunnel creation: allocate_and_dispatch(), allocate_ports_for_container(), CloudflareTunnelManager
Scenarios implemented (playbook-driven)

create-nginx
Creates a container in a new pod (nginx:stable-alpine by default), requests port 80 with use_dns: true.
Waits for allocation then running.
Discovers ingress from pod.port_forwardings (dns_record_hostname or external_ip_address:external_port).
Probes DNS resolution, TCP connectivity, and HTTP GET / expecting 200 and “nginx”.
Teardown deletes the container, then validates:
TCP connectivity to previously known ingress fails
Port forwarding entries for the container are removed from the pod
Code: _scenario_create_container()
add-second-container
Adds a container to the pod created above, waits running, and probes the new container.
Also probes original container still serves traffic.
Deletes only the newly added container.
Code: _scenario_add_to_pod()
negative-image
Uses a bogus image (thisdoesnotexist.invalid:never), expects final status launch_failed.
Validates no ingress is assigned.
Cleans up the failed container record.
Code: _scenario_negative_image()
Runner capabilities

DNS checks are enforced based on playbook defaults; uses DNS if available and also probes external IP:port as fallback unless dns_required is strictly enforced.
Robust waiters with polling and timeouts: allocation, running, deletion.
Data-plane probes:
DNS resolution: dns_resolve()
TCP connect: tcp_connect()
HTTP GET with redirects and content assertions: http_get()
JSON report with timings and per-endpoint assertions written to configurable path.
How to run

Set env vars or edit the playbook:
API_BASE_URL (default http://172.17.0.1:5000/api)
API_KEY
VDC_ID (target UUID)
DNS_REQUIRED=true (enforced in example playbook)
Optional: WEBSOCKET_SERVER_URL for documentation purposes
Review and update smoke/playbook.example.yaml.

Execute:

python smoke/runner.py --playbook smoke/playbook.example.yaml --report smoke/out/report.json
The runner prints summary and writes detailed JSON results.
Notes and extension points

Container restart validations can leverage IaaSClient.container_lifecycle() and assert data-plane blip/resume while worker transitions via container_lifecycle_action().
Additional coverage can be added for volumes, networks, VMs, and pod lifecycle actions reusing the same framework.
The framework stays external: only public API calls and real-world network probes; no internal DB calls.
This delivers the initial smoke test infrastructure and the required test coverage: create/delete container with DNS and data-plane verification, add container to an existing pod with verification for both, and a negative “dodgy image” test that must fail correctly.
2025-08-25 14:20:19 +09:30
coryHawkvelt 08a6c5eae3 Consolidation implemented with a single, phased workflow and small, composable helpers
What changed

Unified two-phase workflow across all scenarios (new pod, add to existing pod, and waiting-host retry):

Phase A (Persist): persist_pod_and_containers()
Phase B (Allocate+Dispatch): allocate_and_dispatch()
Celery task updated to always use the unified flow:

process_workload() now calls persist_pod_and_containers() then allocate_and_dispatch() for container requests.
Deprecated legacy flows replaced with thin forwarding wrappers (kept temporarily for compatibility):

_create_new_pod_flow() forwards to persist+allocate.
_add_to_existing_pod_flow() forwards to persist+allocate.
_build_pod_objects_and_enqueue() forwards to persist+allocate.
allocate_and_dispatch refactored to be safe and idempotent for mixed pods:

Only processes pending-allocation containers via _pending_allocation_containers(), preventing status transition errors and duplicate PF/DNS when appending to a running pod.
Extracted small helpers to keep logic simple and testable:
Host selection: select_host_for_pod()
Waiting-host backoff/jitter: set_waiting_host_and_next_retry()
NS controller alignment and launch params: ensure_nscontroller_on_host()
Port allocation and PF creation: allocate_ports_for_container()
Cloudflare tunnel and DNS: ensure_tunnel_and_dns()
External port selection (retained): find_available_external_port()
Status bulk update (retained): _set_container_statuses()
Cloudflare config standardized:

All references now use CLOUDFLARE_API_TOKEN, CLOUDFLARE_ACCOUNT_ID, CLOUDFLARE_ZONE_ID in allocate_and_dispatch() and ensure_tunnel_and_dns().
Why this satisfies the goal

Single workflow for all container creations and management:

New pod creation and adding to an existing pod both start by persisting to DB and then run allocation+dispatch through the same code path allocate_and_dispatch().
Celery retry for waiting-host continues calling allocate_and_dispatch() which now encapsulates selection, PF/DNS, and dispatch.
Safe status transitions and idempotency:

Only new/pending containers are set to allocated on success; existing running containers are untouched, preventing validation errors in Workload._validate_status_change().
Port mappings and DNS are only created for the new/pending containers to avoid duplicates.
Notes and next steps (tracked in TODOs)

Optional further refinement: extract dispatch and status-setting blocks into helpers for even tighter separation.
Add tests covering new pod, add-to-existing, waiting-host retry, and idempotent re-dispatch behavior.
Update docs and curl samples to show both payload shapes.
Key entry points to review

Persist phase: persist_pod_and_containers()
Allocate+Dispatch: allocate_and_dispatch()
Pod payload builder: build_pod_payload()
Celery task: process_workload()
Retry beat: retry_waiting_host_pods()
This achieves a consolidated, maintainable workflow without introducing a single large function, using small helpers to keep concerns separated and the codebase easier to evol
2025-08-25 12:52:51 +09:30
coryHawkvelt 8d2e79a173 New container for changin content of index.html 2025-08-25 12:02:07 +09:30
coryHawkvelt ee8429af1f add pod status to json output 2025-08-25 10:37:22 +09:30
coryHawkvelt 3d5da8525b Moved celery job 2025-08-25 10:37:14 +09:30
coryHawkvelt 37aa2f2ded Updated logger.py to provide meaningful request IDs outside of Flask.
What changed:

Added a context-based ID source using Python ContextVar, used when explicitly set.
Preserved Flask behavior (reads g.request_id when a request context exists).
Added Celery awareness: when running in a Celery worker/task, it uses the current task’s id (or correlation_id) automatically.
Final fallback now includes process/thread info instead of “no-context”.
Resolution of the reported issue:

Your Celery worker logs will now include RequestID set to the Celery task id (or correlation_id). If not in a Celery task and not in Flask, the logger will emit a structured fallback like pid:1234|thr:MainThread rather than no-context.
For non-Flask, non-Celery jobs:

At the beginning of your job or script, set a correlation/request ID via the logger module’s context setter and clear it when done. This ensures all logs in that execution path include your chosen ID.
No changes required to your existing Flask code paths, and the formatter/handlers remain intact, so log output format is unchanged aside from improved request ID values
2025-08-25 10:37:03 +09:30
coryHawkvelt c8c011c6cf Container pod workflow - Retry if scheduling failed 2025-08-25 09:07:30 +09:30
coryHawkvelt c2cbcc1222 Celery broker url 2025-08-25 05:38:25 +09:30
coryHawkvelt 519a258eda Fully AI split f workload creation and ASAP to database logic - NOT YET REVIEWED 2025-08-25 05:16:30 +09:30
coryHawkvelt 14e2389f74 Add network status field and migrations to suit 2025-08-25 04:18:02 +09:30
coryHawkvelt ef013d8607 Added audit entries to all POST\PUT and DELETE 2025-08-25 00:20:20 +09:30
coryHawkvelt e216953fac Kilocode nicely found all non-standard responses for us! 2025-08-24 23:36:03 +09:30
coryHawkvelt e1fa8de4dc Merge pull request 'upd' (#27) from release into master
Reviewed-on: coryHawkvelt/theapi#27
2025-08-22 15:48:19 +00:00
coryHawkvelt 510edb85c9 Merge pull request 'release' (#26) from release into master
Reviewed-on: coryHawkvelt/theapi#26
2025-08-22 15:46:50 +00:00
coryHawkvelt 5663afea04 Merge pull request 'ssl-ca' (#25) from ssl-ca into master
Reviewed-on: coryHawkvelt/theapi#25
2025-08-22 03:47:55 +00:00
coryHawkvelt beae6b181c moving everything into standard api responses 2025-08-04 07:38:40 +09:30
coryHawkvelt a709b59214 refactor: Simplify host downtime handling with delayed task scheduling 2025-08-03 14:29:02 +09:30
coryHawkvelt f40d8c04e0 feat: Add host downtime monitoring task to handle offline hosts and update workload statuses 2025-08-03 14:21:40 +09:30
coryHawkvelt e78c690039 OpenAI compat almost achived! 2025-08-03 14:09:42 +09:30
coryHawkvelt 8e65534856 minor tidy ups 2025-08-03 12:43:10 +09:30
coryHawkvelt a826205381 refactor: Update delete_container_workload and delete_pod to use standard API response 2025-08-03 12:23:54 +09:30
coryHawkvelt 77d8fab7c7 refactor: Update get_pods and get_pod to use standard api_response pattern 2025-08-03 12:23:15 +09:30
coryHawkvelt e4b3f1d5a2 refactor: Update get_container_workloads to use standard api_response pattern 2025-08-03 12:22:15 +09:30
coryHawkvelt 2284261a71 refactor: Update get_container_workload to use standard api_response pattern 2025-08-03 12:21:39 +09:30
coryHawkvelt 04be1fc190 refactor: Update update_container_workload to use standard api_response 2025-08-03 12:20:55 +09:30
coryHawkvelt bd0cb4836a refactor: Move import statement to top of file for better organization 2025-08-03 12:20:51 +09:30
coryHawkvelt cbd675ed73 emoved blacklisting, it broke the logic and the fake status update was ugly and duplicated code.. will come back to this later when we do the reconciliation 2025-08-03 12:04:18 +09:30
coryHawkvelt db1c41426d feat: Add container status reconciliation and force status update methods 2025-08-03 11:42:18 +09:30
coryHawkvelt 8be65b95ba refactor: Update ContainerTask to accept docker_monitor in constructor 2025-08-03 11:42:15 +09:30
coryHawkvelt d9877bdd94 refactor: Add Docker monitor blacklisting to container lifecycle operations 2025-08-03 11:25:12 +09:30
coryHawkvelt 4e1de0f37e fix: Improve container start and recreation handling in Docker worker 2025-08-03 11:00:45 +09:30