Commit Graph
100 Commits
Author SHA1 Message Date
coryHawkvelt c42c352a78 NOT REVIEWED - Failover logic enhancements 2025-11-06 13:28:12 +10:30
coryHawkvelt 042fdb8478 docs and notes 2025-11-06 13:25:40 +10:30
coryHawkvelt 774e5c190b refactor(worker): defer manager initialization to worker processes
move multiprocessing manager initialization from constructor to run() method in both DockerMonitor and LibvirtMonitor to avoid issues with forking managers across processes

fix(worker): prevent stop event errors during shutdown

add null checks for stop_event and manager before attempting to use or shutdown them in monitor classes

refactor(worker): simplify process management in main

replace processes list with direct worker_proc handling for cleaner shutdown logic

feat(worker): add error handling for monitor process startup

wrap monitor process creation in try-except blocks to catch and log startup failures
2025-11-04 15:28:45 +10:30
coryHawkvelt 7fd7cd9aaa feat(worker): implement monitor creation inside worker process and auto-restart
- Move DockerMonitor and LibvirtMonitor instantiation into the worker process after successful join
- Add support for disabling monitors via no_docker and no_libvirt flags
- Implement automatic restart of worker process on unexpected termination
- Add reconnection logic to WorkerClient with exponential backoff
- Ensure container reconciliation runs on both initial join and reconnections
2025-11-04 15:12:09 +10:30
coryHawkvelt 87adb4938b feat(tasks): add offline host check with backoff in allocation retry
Introduce handling for pods with assigned but offline hosts in the retry task.
Increment allocation attempts, apply exponential backoff with jitter, and
schedule next retry accordingly. Skip immediate reallocation if host is offline
to prevent unnecessary attempts. Configuration values are used for base delay,
backoff factor, jitter, and max delay.
2025-11-04 14:41:59 +10:30
coryHawkvelt 8586ace1d6 refactor(container): optimize port allocation and NSController handling for migration
Refactored the container allocation logic to improve migration support:
- Moved NSController handling directly into main allocation flow
- Added migration-aware port allocation with existing port reuse
- Simplified function signatures by removing ns_lp parameter dependency
- Moved DB commit earlier in the process for better transaction management
- Improved stale port mapping cleanup with proper error handling
- Fixed typos in logging messages
2025-11-04 14:02:56 +10:30
coryHawkvelt 1027f77011 feat(worker): add full status sync, map statuses
- Add send_full_statuses_for_host to emit a full snapshot after
  reconciliation for convergence. Supports "host" or "pod" scope via
  WORKER_FULL_STATUS_SCOPE, filters by managed_by label, and skips
  neutral states.
- Map Docker container statuses to websocket_server event types in
  WorkerClient: "running" -> running, "exited"/"dead"/"stopped" -> die,
  skip others. Align details.status and Action accordingly to avoid
  unsupported events.
- Abort launching regular containers if NSController launch fails and
  improve logging/injection behavior.

These changes improve state convergence and robustness while preventing
invalid event emissions.
2025-11-04 11:58:53 +10:30
coryHawkvelt 77140d1b95 fix(pod): reassign host if mismatched in allocation
- Update host assignment condition to reassign if workload host ID is set but does not match the selected host
- Add debug logging for host selection and pre-populated host cases
- Add TODO for setting pod state to allocated after container allocation
2025-11-03 12:42:55 +10:30
coryHawkvelt bbfab8b8b7 docs and examples 2025-10-27 08:42:52 +10:30
coryHawkvelt 0237024d1c feat(task-assigner): add debug flag for task assignment logging
Introduce `ENABLE_TASK_ASSIGNMENT_DEBUG` configuration to control verbose
logging in task assignment, Redis subscription, and worker management flows.
This reduces log noise by default while allowing detailed debugging when
enabled.
2025-10-27 08:42:45 +10:30
coryHawkvelt 54ac191f65 Merge branches 'master' and 'master' of ssh://source.hawkless.id.au:222/coryHawkvelt/theapi 2025-10-27 08:00:56 +10:30
coryHawkvelt 3cd200477e feat(worker): add OVS bridge scanner task
Integrate OVSBridgeScannerTask into WorkerClient to scan for local OVS
bridges using ovs-vsctl, report them to the API endpoint, and handle
initialization with error logging. Includes retry logic for API calls
and graceful fallback if OVS is unavailable.
2025-10-27 07:59:03 +10:30
coryHawkvelt 3070b0c703 feat(api): add endpoints for retrieving and updating OVS bridges
Add GET endpoint to fetch non-deleted OVS bridge names for a workload host.

Add POST endpoint to update OVS bridges: validate input, soft-delete removed
bridges, add new bridges, update timestamps for unchanged ones, with error
handling and logging.
2025-10-27 07:46:08 +10:30
coryHawkvelt 67df40def2 feat(container): add lifecycle state management and rename cpu_shares to cpu
Introduce support for container states (present, absent, started, stopped,
restarted) via xcloudify_containers_state variable in Ansible role.

Update container specs and module to use 'cpu' instead of 'cpu_shares' for
allocation, with validation and API adjustments.

Update examples and tasks for state handling, debug output, and error conditions.

BREAKING CHANGE: 'cpu_shares' renamed to 'cpu' in container configurations and module options.
2025-10-27 07:26:43 +10:30
coryHawkvelt 3bca326b51 feat(websocket): add debug flag for ping operations
Introduce `ENABLE_WEBSOCKET_PING_DEBUG` environment variable to
conditionally enable verbose logging for websocket ping sends and
pong receptions, reducing log noise in production.
2025-10-27 07:21:16 +10:30
coryHawkvelt 05f9862f16 refactor(container): fetch pod via query using pod_id 2025-10-27 07:02:34 +10:30
coryHawkvelt c9d0a1cef7 chore(container): remove commented-out pod mapping code 2025-10-27 06:52:34 +10:30
coryHawkvelt 309b132e9c feat(scheduling): add generic resource filter with overcommit support
Refactor HasSufficientResources to support arbitrary resource types via
host.get_resource_utilization(), applying per-resource overcommit ratios
(default 1.0). Introduce dynamic aggregation of pod resource requirements
from WorkloadResourceUsage database entries, replacing hardcoded minima.

Integrate the filter early in select_host_for_pod pipeline when CPU/RAM
requirements exceed zero, preserving filter order for efficiency.

Adjust global cloudflared sidecar limits: CPU to 1 (from 3), RAM to 64MiB
(from 132MiB) for reduced overhead.
2025-10-27 05:57:44 +10:30
coryHawkvelt 4a3681c667 feat(cloudflare): add reconciliation worker and beat task
- Introduce CloudflareReconciliationWorker to reconcile Cloudflare
  tunnels/DNS with DB and clean up stale resources
- Add Celery task tasks.cloudflare_reconciliation and schedule it
  every 5 minutes via beat
- Provide standalone script to run the reconciliation manually
- Add optional scheduler helper for periodic runs

Also:
- Update tunnel lookup to not ignore tunnels in down status so
  cleanup can find and delete them

No breaking changes.
2025-10-27 04:16:11 +10:30
coryHawkvelt 2d9ee4882d fix(container-deletion): pass workload object to cloudflare cleanup function
- modify `cleanup_cloudflare_resources` to accept workload object instead of container id
- remove redundant workload lookup inside the function
- update usage in `process_container_deletion` to pass container object
- adjust resource usage assignments in `ensure_tunnel_and_dns` to use sidecar id
- remove obsolete `_cleanup_nscontroller_cloudflare_resources` function
2025-10-27 03:33:04 +10:30
coryHawkvelt ff0a235355 feat(workload): introduce global resource limits for cloudflared sidecar
refactor resource limit definitions by using global variables for cpu and memory, and track resource usage in WorkloadResourceUsage entries for the NSController workload
2025-10-27 03:12:47 +10:30
coryHawkvelt 23bddecd21 refactor(deletion): consolidate deletion logic into unified handler
Move all workload and pod deletion logic into a new unified deletion handler module. This includes marking workloads as pending-deleted, cleaning up associated resources (port forwardings, volume mappings, network ports, Cloudflare DNS records and tunnels), and updating pod statuses after container deletions. The change reduces code duplication across multiple files and ensures consistent cleanup behavior.

The new app/utils/deletion_handler.py module contains functions for:
- mark_workload_pending_deleted: marks workloads and associated resources as pending-deleted
- cleanup_workload_resources: performs actual resource deletion
- cleanup_cloudflare_resources_for_workload: handles Cloudflare-specific cleanup
- mark_pod_pending_deleted: marks all containers in a pod for deletion
- update_pod_status_after_deletion: updates pod status based on container states

Modified existing code to use these new consolidated functions instead of repeating similar logic in multiple places.
2025-10-27 01:21:37 +10:30
coryHawkvelt 0aa048ba02 junk 2025-10-27 00:00:53 +10:30
coryHawkvelt ea1c3a6e82 refactor(error-handling): introduce custom exceptions for cloudflare and port errors
Move NoFreePortsError to dedicated exceptions module and add CloudFlareException for
Cloudflare-specific failures. Update ensure_tunnel_and_dns to raise CloudFlareException
on errors. Integrate set_waiting_host_and_next_retry in allocation_retry and
process_workload_request tasks for consistent failure handling with exponential backoff.
Handle pod ID as string in set_waiting_host_and_next_retry and add logging.
2025-10-26 23:53:20 +10:30
coryHawkvelt eb4243b97c Dont override cpu and meory for sidecards, it's explicitly defined when creating them 2025-10-26 23:49:54 +10:30
coryHawkvelt e31fda6220 feat(models): add resource utilization methods for hosts and workloads
Add get_resource_utilization to WorkloadHost to compute available and used
resources dynamically from pooled resources and active workloads.

Add get_resource_usage to Workload to return list of used resources for the
workload.

Fix minor typo in TODO comment and add SQLAlchemy func import.
2025-10-26 23:11:05 +10:30
coryHawkvelt 0989a66198 refactor(container): rename cpu_shares to cpu and update resource defaults
Renamed the CPU resource parameter from 'cpu_shares' to 'cpu' across API
validation, payload building, resource usage tracking, and worker tasks for
consistency and simplicity. Adjusted default mem_limit from 128MB to 129MB.
Introduced global constants for NSController resources (5 CPU, 199MB) and set
tunnel sidecar limits (3 CPU, 257MB). Removed unused import.

BREAKING CHANGE: Container API payloads now expect 'cpu' key instead of 'cpu_shares
2025-10-26 22:53:14 +10:30
coryHawkvelt a15f43d03b feat(api): add host utilization endpoint and enrich responses
Introduce _enrich_host_data helper to integrate resource utilization data
into host responses, replacing manual resource queries.

- Update get_workload_host and get_workload_hosts to use the helper for
  pooled resources based on utilization
- Enrich active workloads with get_resource_usage() in get_all_active_workloads
- Change resource_type from "memory" to "ram" during host enrollment
- Add GET /workload_hosts/<id>/utilization endpoint with UUID validation
  and error handling
2025-10-26 22:52:22 +10:30
coryHawkvelt a9e25fe6a7 refactor(worker): use configurable WORKER_ID for container management
Replace hardcoded "worker_agent" string with dynamic WORKER_ID from settings
across Docker monitoring, client operations, and container tasks. This enables
unique identification for multi-worker environments, preventing container
management conflicts.
2025-10-26 21:08:32 +10:30
coryHawkvelt f8721a49ec minor tidy 2025-10-26 21:08:21 +10:30
coryHawkvelt 9fa32731be ansible role 2025-09-17 20:27:50 +10:00
coryHawkvelt 9df4c965d3 upd todo 2025-08-27 17:28:29 +09:30
coryHawkvelt f115bf05a5 Dont duplicate containers at lunch time 2025-08-27 17:28:22 +09:30
coryHawkvelt a2596b954a default all containers to 1gb a cpu 2025-08-27 17:27:10 +09:30
coryHawkvelt 6005c61cc7 Remove workload_host_id from WorkloadResourceUsage 2025-08-27 16:45:51 +09:30
coryHawkvelt fac6d55ff0 Smoke test
Implemented a playbook-driven external smoke test framework and initial scenarios with DNS-inclusive checks.

What was added/changed

New files:
smoke/config.py – playbook schema, env overrides, loader, scenario filtering
smoke/runner.py – runner with three scenarios, waiters, DNS/TCP/HTTP probes, JSON reporting
smoke/playbook.example.yaml – sample playbook using IP-based defaults with DNS checks
smoke/README.md – usage and operation docs
smoke/init.py – package marker
API client uplift:
Added IaaSClient.container_lifecycle() to call POST /workloads/containers/<id>/lifecycle/<action> for future restart tests
Leveraged existing client methods:
IaaSClient.get_container() and IaaSClient.delete_container()
IaaSClient.get_container_pods() and IaaSClient.get_container_pod()
Server routes/flow relied on:
Create (async, DB-first): /workloads/containers
Worker status update: update_container_workload()
Pod introspection with DNS and port_forwardings: get_pods(), get_pod()
Allocation + DNS/Tunnel creation: allocate_and_dispatch(), allocate_ports_for_container(), CloudflareTunnelManager
Scenarios implemented (playbook-driven)

create-nginx
Creates a container in a new pod (nginx:stable-alpine by default), requests port 80 with use_dns: true.
Waits for allocation then running.
Discovers ingress from pod.port_forwardings (dns_record_hostname or external_ip_address:external_port).
Probes DNS resolution, TCP connectivity, and HTTP GET / expecting 200 and “nginx”.
Teardown deletes the container, then validates:
TCP connectivity to previously known ingress fails
Port forwarding entries for the container are removed from the pod
Code: _scenario_create_container()
add-second-container
Adds a container to the pod created above, waits running, and probes the new container.
Also probes original container still serves traffic.
Deletes only the newly added container.
Code: _scenario_add_to_pod()
negative-image
Uses a bogus image (thisdoesnotexist.invalid:never), expects final status launch_failed.
Validates no ingress is assigned.
Cleans up the failed container record.
Code: _scenario_negative_image()
Runner capabilities

DNS checks are enforced based on playbook defaults; uses DNS if available and also probes external IP:port as fallback unless dns_required is strictly enforced.
Robust waiters with polling and timeouts: allocation, running, deletion.
Data-plane probes:
DNS resolution: dns_resolve()
TCP connect: tcp_connect()
HTTP GET with redirects and content assertions: http_get()
JSON report with timings and per-endpoint assertions written to configurable path.
How to run

Set env vars or edit the playbook:
API_BASE_URL (default http://172.17.0.1:5000/api)
API_KEY
VDC_ID (target UUID)
DNS_REQUIRED=true (enforced in example playbook)
Optional: WEBSOCKET_SERVER_URL for documentation purposes
Review and update smoke/playbook.example.yaml.

Execute:

python smoke/runner.py --playbook smoke/playbook.example.yaml --report smoke/out/report.json
The runner prints summary and writes detailed JSON results.
Notes and extension points

Container restart validations can leverage IaaSClient.container_lifecycle() and assert data-plane blip/resume while worker transitions via container_lifecycle_action().
Additional coverage can be added for volumes, networks, VMs, and pod lifecycle actions reusing the same framework.
The framework stays external: only public API calls and real-world network probes; no internal DB calls.
This delivers the initial smoke test infrastructure and the required test coverage: create/delete container with DNS and data-plane verification, add container to an existing pod with verification for both, and a negative “dodgy image” test that must fail correctly.
2025-08-25 14:20:19 +09:30
coryHawkvelt 08a6c5eae3 Consolidation implemented with a single, phased workflow and small, composable helpers
What changed

Unified two-phase workflow across all scenarios (new pod, add to existing pod, and waiting-host retry):

Phase A (Persist): persist_pod_and_containers()
Phase B (Allocate+Dispatch): allocate_and_dispatch()
Celery task updated to always use the unified flow:

process_workload() now calls persist_pod_and_containers() then allocate_and_dispatch() for container requests.
Deprecated legacy flows replaced with thin forwarding wrappers (kept temporarily for compatibility):

_create_new_pod_flow() forwards to persist+allocate.
_add_to_existing_pod_flow() forwards to persist+allocate.
_build_pod_objects_and_enqueue() forwards to persist+allocate.
allocate_and_dispatch refactored to be safe and idempotent for mixed pods:

Only processes pending-allocation containers via _pending_allocation_containers(), preventing status transition errors and duplicate PF/DNS when appending to a running pod.
Extracted small helpers to keep logic simple and testable:
Host selection: select_host_for_pod()
Waiting-host backoff/jitter: set_waiting_host_and_next_retry()
NS controller alignment and launch params: ensure_nscontroller_on_host()
Port allocation and PF creation: allocate_ports_for_container()
Cloudflare tunnel and DNS: ensure_tunnel_and_dns()
External port selection (retained): find_available_external_port()
Status bulk update (retained): _set_container_statuses()
Cloudflare config standardized:

All references now use CLOUDFLARE_API_TOKEN, CLOUDFLARE_ACCOUNT_ID, CLOUDFLARE_ZONE_ID in allocate_and_dispatch() and ensure_tunnel_and_dns().
Why this satisfies the goal

Single workflow for all container creations and management:

New pod creation and adding to an existing pod both start by persisting to DB and then run allocation+dispatch through the same code path allocate_and_dispatch().
Celery retry for waiting-host continues calling allocate_and_dispatch() which now encapsulates selection, PF/DNS, and dispatch.
Safe status transitions and idempotency:

Only new/pending containers are set to allocated on success; existing running containers are untouched, preventing validation errors in Workload._validate_status_change().
Port mappings and DNS are only created for the new/pending containers to avoid duplicates.
Notes and next steps (tracked in TODOs)

Optional further refinement: extract dispatch and status-setting blocks into helpers for even tighter separation.
Add tests covering new pod, add-to-existing, waiting-host retry, and idempotent re-dispatch behavior.
Update docs and curl samples to show both payload shapes.
Key entry points to review

Persist phase: persist_pod_and_containers()
Allocate+Dispatch: allocate_and_dispatch()
Pod payload builder: build_pod_payload()
Celery task: process_workload()
Retry beat: retry_waiting_host_pods()
This achieves a consolidated, maintainable workflow without introducing a single large function, using small helpers to keep concerns separated and the codebase easier to evol
2025-08-25 12:52:51 +09:30
coryHawkvelt 8d2e79a173 New container for changin content of index.html 2025-08-25 12:02:07 +09:30
coryHawkvelt ee8429af1f add pod status to json output 2025-08-25 10:37:22 +09:30
coryHawkvelt 3d5da8525b Moved celery job 2025-08-25 10:37:14 +09:30
coryHawkvelt 37aa2f2ded Updated logger.py to provide meaningful request IDs outside of Flask.
What changed:

Added a context-based ID source using Python ContextVar, used when explicitly set.
Preserved Flask behavior (reads g.request_id when a request context exists).
Added Celery awareness: when running in a Celery worker/task, it uses the current task’s id (or correlation_id) automatically.
Final fallback now includes process/thread info instead of “no-context”.
Resolution of the reported issue:

Your Celery worker logs will now include RequestID set to the Celery task id (or correlation_id). If not in a Celery task and not in Flask, the logger will emit a structured fallback like pid:1234|thr:MainThread rather than no-context.
For non-Flask, non-Celery jobs:

At the beginning of your job or script, set a correlation/request ID via the logger module’s context setter and clear it when done. This ensures all logs in that execution path include your chosen ID.
No changes required to your existing Flask code paths, and the formatter/handlers remain intact, so log output format is unchanged aside from improved request ID values
2025-08-25 10:37:03 +09:30
coryHawkvelt c8c011c6cf Container pod workflow - Retry if scheduling failed 2025-08-25 09:07:30 +09:30
coryHawkvelt c2cbcc1222 Celery broker url 2025-08-25 05:38:25 +09:30
coryHawkvelt 519a258eda Fully AI split f workload creation and ASAP to database logic - NOT YET REVIEWED 2025-08-25 05:16:30 +09:30
coryHawkvelt 14e2389f74 Add network status field and migrations to suit 2025-08-25 04:18:02 +09:30
coryHawkvelt ef013d8607 Added audit entries to all POST\PUT and DELETE 2025-08-25 00:20:20 +09:30
coryHawkvelt e216953fac Kilocode nicely found all non-standard responses for us! 2025-08-24 23:36:03 +09:30
coryHawkvelt e1fa8de4dc Merge pull request 'upd' (#27) from release into master
Reviewed-on: coryHawkvelt/theapi#27
2025-08-22 15:48:19 +00:00
coryHawkvelt 510edb85c9 Merge pull request 'release' (#26) from release into master
Reviewed-on: coryHawkvelt/theapi#26
2025-08-22 15:46:50 +00:00
coryHawkvelt 5663afea04 Merge pull request 'ssl-ca' (#25) from ssl-ca into master
Reviewed-on: coryHawkvelt/theapi#25
2025-08-22 03:47:55 +00:00
coryHawkvelt beae6b181c moving everything into standard api responses 2025-08-04 07:38:40 +09:30
coryHawkvelt a709b59214 refactor: Simplify host downtime handling with delayed task scheduling 2025-08-03 14:29:02 +09:30
coryHawkvelt f40d8c04e0 feat: Add host downtime monitoring task to handle offline hosts and update workload statuses 2025-08-03 14:21:40 +09:30
coryHawkvelt e78c690039 OpenAI compat almost achived! 2025-08-03 14:09:42 +09:30
coryHawkvelt 8e65534856 minor tidy ups 2025-08-03 12:43:10 +09:30
coryHawkvelt a826205381 refactor: Update delete_container_workload and delete_pod to use standard API response 2025-08-03 12:23:54 +09:30
coryHawkvelt 77d8fab7c7 refactor: Update get_pods and get_pod to use standard api_response pattern 2025-08-03 12:23:15 +09:30
coryHawkvelt e4b3f1d5a2 refactor: Update get_container_workloads to use standard api_response pattern 2025-08-03 12:22:15 +09:30
coryHawkvelt 2284261a71 refactor: Update get_container_workload to use standard api_response pattern 2025-08-03 12:21:39 +09:30
coryHawkvelt 04be1fc190 refactor: Update update_container_workload to use standard api_response 2025-08-03 12:20:55 +09:30
coryHawkvelt bd0cb4836a refactor: Move import statement to top of file for better organization 2025-08-03 12:20:51 +09:30
coryHawkvelt cbd675ed73 emoved blacklisting, it broke the logic and the fake status update was ugly and duplicated code.. will come back to this later when we do the reconciliation 2025-08-03 12:04:18 +09:30
coryHawkvelt db1c41426d feat: Add container status reconciliation and force status update methods 2025-08-03 11:42:18 +09:30
coryHawkvelt 8be65b95ba refactor: Update ContainerTask to accept docker_monitor in constructor 2025-08-03 11:42:15 +09:30
coryHawkvelt d9877bdd94 refactor: Add Docker monitor blacklisting to container lifecycle operations 2025-08-03 11:25:12 +09:30
coryHawkvelt 4e1de0f37e fix: Improve container start and recreation handling in Docker worker 2025-08-03 11:00:45 +09:30
coryHawkvelt 506ecdf038 fix: Improve container lifecycle error handling and reporting 2025-08-03 10:57:30 +09:30
coryHawkvelt 60a388754f feat: Add support for container lifecycle events (start, stop, restart) 2025-08-03 10:51:58 +09:30
coryHawkvelt 6a9bec1e7e feat: Add container and pod lifecycle management APIs 2025-08-03 10:41:47 +09:30
coryHawkvelt 5ddf95af29 refactor: Remove redundant import and simplify AuditEntry logging 2025-08-03 10:41:38 +09:30
coryHawkvelt 1731e2c93a feat: Expand workload host endpoints to include pooled and fixed resources 2025-08-03 10:00:27 +09:30
coryHawkvelt 95b52a70b6 feat: Add CPU and RAM capture during workload host enrollment 2025-08-03 09:54:08 +09:30
coryHawkvelt 2fa3dc3fe6 feat: Add region enrollment key validation for workload host enrollment 2025-08-03 09:54:05 +09:30
coryHawkvelt c081d5046d feat: Add support for NFS volume processing with SHAREDFS_ROOT
This commit implements the following changes:
- Add SHAREDFS_ROOT configuration option with default value "/mnt/shared"
- Create _process_storage_volumes method to handle NFS volume directory creation
- Update container launch and validation to support NFS volume mounting
- Ensure storage volumes are processed before container creation in pod update
2025-08-03 09:47:15 +09:30
coryHawkvelt c183afa9b3 refactor: Delete volume mappings without removing volumes during container deletion 2025-08-03 09:30:41 +09:30
coryHawkvelt 03f4781349 fix: Normalize CPU shares to Docker's valid range to prevent container launch errors 2025-08-03 09:26:21 +09:30
coryHawkvelt 7518583b00 feat: Implement CPU shares and memory limit support for container creation 2025-08-03 09:19:22 +09:30
coryHawkvelt 06ab7db08b feat: Add WorkloadResourceUsage model to track resource usage per workload 2025-08-03 09:14:25 +09:30
coryHawkvelt 9a50176150 refactor: Add CPU and memory limits for containers with default sidecar resources 2025-08-03 09:08:28 +09:30
coryHawkvelt 7738ed5a0a refactor: Change volume handling from volumes to storage in pod payload 2025-08-03 09:08:26 +09:30
coryHawkvelt a94b84103f feat: Add support for storage volumes in container workload creation 2025-08-03 08:42:48 +09:30
coryHawkvelt 1e8d652496 refactor: Organize imports in create_workload_container.py 2025-08-03 08:42:45 +09:30
coryHawkvelt a005f9dcc9 feat: Add endpoints to list volumes for workloads and workloads for volumes 2025-08-03 06:58:08 +09:30
coryHawkvelt adc08c1957 feat: Add comprehensive volume API routes with validation and standard responses 2025-08-03 06:56:33 +09:30
coryHawkvelt 3155e7791c feat: Add audit log routes for containers, pods, and action types
This commit introduces new API endpoints for retrieving audit logs:
- `/api/audit/container/<container_id>` to get container-specific logs
- `/api/audit/pod/<pod_id>` to get pod-specific logs
- `/api/audit/actions` to list available audit action types

The implementation includes:
- Comprehensive error handling
- Pagination support
- Filtering by action type
- Standard API response formatting
- Detailed logging for debugging
2025-08-03 05:16:31 +09:30
coryHawkvelt 09639ddc10 feat: Add comprehensive audit logging for container creation process 2025-08-03 05:11:37 +09:30
coryHawkvelt a7c70165bd documentation, tify up and vnc token into vm route 2025-07-28 05:08:49 +09:30
coryHawkvelt a455e32626 adjust allowed state transitions 2025-07-28 00:53:55 +09:30
coryHawkvelt 13ff5159bb supress debug logs 2025-07-28 00:53:44 +09:30
coryHawkvelt cc1e2c7f84 fixed redis subscription keysapce issue mismatch 2025-07-28 00:53:37 +09:30
coryHawkvelt c40f6aeef7 pplaying with scheduler filters 2025-07-28 00:53:08 +09:30
coryHawkvelt f40b0543cf dummied the ovs response 2025-07-28 00:52:55 +09:30
coryHawkvelt eb6133590f Configured worker for docker 2025-07-28 00:52:39 +09:30
coryHawkvelt 26921cab1e added a sidecar flag 2025-07-26 10:51:51 +09:30
coryHawkvelt 904b2e011d Adding a new container to an existing pod doesnt break everything! 2025-07-26 10:31:59 +09:30
coryHawkvelt 6e21e49681 Tuneing ping timeouts 2025-07-26 04:00:46 +09:30
coryHawkvelt 45acc4eec6 Tidy up dead files 2025-07-26 03:01:21 +09:30
coryHawkvelt 93d4a96226 Added offline->online transition for workers that were erronously marked as offline 2025-07-26 02:58:57 +09:30
coryHawkvelt 070429ef46 Implimented a ping\pong monitoring system work worker liveness. Will handle the situation where a websocket server crashes.
The only thing this doesnt handle is when a client goes into offline due to a transient link failure but it comes back online. We want to avoid updating the DB every time a pong is RX'd.. Perhaps we do a seocndary corss checks with offline hosts?
2025-07-26 02:52:56 +09:30
coryHawkvelt 7613ce4be7 docs and sandbox update 2025-07-26 00:54:28 +09:30