Observability and troubleshooting
Monitor PIG as a flow: hosts upload sessions, the worker stores traces, analysis completes, and findings and status synchronize to the Promptless Dashboard. A healthy pod covers only part of that flow.
The analyzer emits structured logs by default. Datadog tracing and Sentry error reporting are optional. The supervisor reports release and migration status for deployments with automatic updates.
Checks to keep visible
Section titled “Checks to keep visible”| Stage | Evidence of progress | Investigate when |
|---|---|---|
| Host collection | Recent successful collector status and host check-in. | An active pilot host stops checking in or uploading. |
| Durable storage | Recent agent_traces.last_ingested_at, a completed canonical object, and a readable object at the recorded location. | Storage errors appear, objects remain pending, or active hosts stop producing stored traces. |
| Analysis | Trace analysis run succeeded, correlated by trace_record_id and analysis_run_id. | Failures repeat or eligible quiet sessions do not complete. |
| Hosted synchronization | Deployment status and newly projected trace metadata or findings reach Promptless. | Local work completes but hosted state remains stale. |
| Remediation | An outcome explaining a pull request, external owner, or missing evidence. | A finding has no visible progress after prerequisites are met. |
Set thresholds around your team’s working hours, expected trace volume, and quiet window. An idle host or a successful analysis with zero findings does not indicate a failure. Establish a baseline with the pilot-session checks before alerting.
Read structured logs
Section titled “Read structured logs”For the deployment name used in the default Helm guide:
kubectl --namespace pig logs deployment/acme-analyzer --since=1hkubectl --namespace pig get podsThe worker writes structured JSON logs to container stdout. Analysis-process logs join the service’s log stream. Your platform’s log collector must forward those logs if you want retention or centralized search; stdout alone does not send them to Datadog.
Correlate analysis messages with analysis_run_id, trace_record_id, analyzer_version, and repository_sha. Successful runs include successful_finding_write_count; failed runs include failure_category. Keep these identifiers when contacting support, along with the PIG release and time window.
If a pod is restarting, inspect its previous logs with kubectl -n pig logs POD_NAME --previous. Avoid dumping Secret contents or full session transcripts into a support ticket.
Optional telemetry integrations
Section titled “Optional telemetry integrations”Forward structured stdout logs with your existing cluster log collector. Apply your organization’s access and retention controls to operational context and any session data collected by optional telemetry.
For the operator-managed worker chart, the manual Helm reference covers Datadog and Sentry settings. Those chart values are not PIGDeployment fields.
Diagnose by the first failing stage
Section titled “Diagnose by the first failing stage”Host cannot enroll or upload
Section titled “Host cannot enroll or upload”Check the host credential, hosted deployment registration, hostname, TLS certificate, and connectivity from the host’s network. Enrollment and upload routes must reach the worker, not just /healthz. Review host enrollment troubleshooting.
Host checks in, but storage does not complete
Section titled “Host checks in, but storage does not complete”Inspect the exact session in agent_traces, including trace_object_status and trace_object_last_error. Check PostgreSQL availability and object-storage authorization, location, prefix, and encryption-key permissions. Keep the same bucket and database while fixing access.
Storage succeeds, but analysis fails
Section titled “Storage succeeds, but analysis fails”Confirm that analysis is enabled, the session is eligible after its quiet window, and the repository and model settings are complete. Use the run’s failure category to narrow the investigation. Test repository read access and model authorization independently; neither follows from successful object-storage access.
The worker retries transient failures during an analysis run. If the run reaches failed, fix the dependency and contact Promptless for supported re-analysis. Preserve its run ID and stored traces. An unchanged failed session revision does not automatically retry after recovery. Avoid ad hoc backfill Jobs or edits to analysis rows.
Local analysis succeeds, but hosted state is stale
Section titled “Local analysis succeeds, but hosted state is stale”Check outbound access to Promptless, the installation credential, and the configuration hash. If the credential was rotated, confirm that the Secret contains the replacement and the analyzer has restarted. Finding projection also requires the hosted repository connection to work. A zero-finding run has no new issue to project. Check the run and deployment status before treating a missing issue as a failure.
An upgrade is blocked or fails
Section titled “An upgrade is blocked or fails”Inspect the current and target releases, conditions, and supervisor logs:
kubectl describe pigdeployment acme --namespace pigkubectl logs deployment/pig-supervisor --namespace pig-system --since=1hkubectl get jobs --namespace pigPreserve the failed migration Job’s logs before retrying. Resolve the reported dependency through its owner: Terraform for cloud resources, your secret manager for credentials, and PIG for application releases. See updates and recovery before choosing a rollback.
For the manual worker chart, inspect job/pig-trace-analyzer-migrate and helm history pig-trace-analyzer --namespace pig. The chart removes successful hook Jobs and replaces the earlier hook on the next attempt.
These names assume the release replacement is complete. For an existing installation, find its names with helm list --namespace pig and kubectl get deployments,jobs --namespace pig. Follow the manual release replacement before installing under the new name.
Preserve data during recovery
Section titled “Preserve data during recovery”Keep PostgreSQL backups and object-storage retention and version recovery under your organization’s recovery policy. Rehearse restoring them as a consistent set: restoring only the database or only objects can leave records and content out of sync.
Restarting the worker can rebuild its disposable repository mirror. It should retain the same registered deployment identity, PostgreSQL database, and trace prefix. Do not clear host collection state, delete traces, or reset database tables to force a retry.
For Helm-managed releases, coordinate restoration and any compatible application rollback with Promptless. For the supervisor’s pause, blocked-upgrade, and recovery behavior, see Manage updates and recovery.