Troubleshooting
Use this guide to triage common Bitaic monitoring failures, run the safest available diagnostic checks, and collect the details support needs when an issue cannot be resolved from the dashboard or CLI.
Identify which signal changed first: certificate, domain, DNS, endpoint availability, alert delivery, or an enrolled private-beta product. That usually points to the fastest fix.
Quick Triage
- Open the affected Bitaic target and note the product area, target name, alert state, last check time, and most recent status change.
- Compare the Bitaic state with the owning system: certificate provider, registrar, DNS provider, service endpoint, or integration destination.
- Review the relevant configuration field, such as
ssl_alert_days,expected_nameservers,expected_values, ormax_latency_ms. - For enrolled private-beta products, use the beta-specific troubleshooting path provided with the workspace.
- Confirm the alert route is enabled and reaches the team responsible for the monitored target.
- If the issue remains unclear, collect the escalation packet below and contact support.
Diagnostic Commands
| Command or check | Use when | Expected signal | Next action |
|---|---|---|---|
bitaic version | The CLI behavior does not match docs or support guidance. | The installed CLI version is visible. | Include the version in the support packet. |
bitaic agent status endpoint | Certificate or Endpoint Monitoring targets are missing, stale, or not checking. | Endpoint agent running state, last check, and config status. | Restart or refresh collection if the agent is stopped or stale. |
bitaic endpoint check <url> | An endpoint is down, degraded, slow, or serving a new HTTPS certificate. | Availability, final status code, redirect count, response time, uptime state, and alert state. | Compare the result with dashboard history and the service owner. |
bitaic dns check <record_name> --type <record_type> | A DNS record is mismatched, slow, unavailable, or still propagating. | Resolver regions, recursive answers, authoritative answer, response code, TTL, DNSSEC state, propagation state, latency, and alert state. | Compare selected resolver regions with authoritative DNS and the configured expected values. |
| Dashboard target history | Domain, DNS, integration, or API details need review. | Last check time, latest state, alert history, and owner route. | Capture the target, timestamp, state, and related configuration. |
Product Troubleshooting Matrix
| Product | Common failures | Quick checks | Escalate when |
|---|---|---|---|
| Certificate Monitoring | Certificate expiration alerts, browser warnings after renewal, or missing HTTPS targets. | Check ssl_alert_days, certificate expiration, served certificate chain, endpoint agent status, and alert route. | Bitaic still shows the old certificate after the renewed certificate is deployed and a fresh check runs. |
| Domain Monitoring | At-risk, degraded, unavailable, or unknown domain state; unexpected renewal, nameserver, DNSSEC, or registrar-lock signal. | Compare the configured domain with registrar records, expected nameservers, DNSSEC expectation, registrar-lock expectation, renewal window, RDAP/WHOIS retrieval state, last check time, and alert history. | The owning registrar state looks healthy but Bitaic continues to show risk or unknown status. |
| DNS Monitoring | Lookup failure, unexpected record answer, propagation drift, or slow resolution. | Review dns_checks, record type, expected_values, alert_on_mismatch, max_latency_ms, resolver regions, authoritative result, expected_ttl_seconds, dnssec_validation, propagation grace, and last check time. | The configured expected answer matches the provider state but Bitaic continues to report mismatch or lookup failure. |
| Endpoint Monitoring | Endpoint down, degraded, slow, missing from dashboard, or alerting to the wrong team. | Run bitaic endpoint check <url>, review the endpoint configuration, compare final status code, redirect count, response-time history, and confirm the configured warning and critical response-time thresholds. | The endpoint is reachable from the expected network path but Bitaic still reports unavailable, degraded, or unknown. |
Health Monitoring and Windows Event Monitoring remain private beta only and release publicly in Q1 2027. Detailed beta troubleshooting stays in beta-specific documentation for enrolled workspaces.
Alert Delivery Checks
- Confirm the monitored target is attached to the expected alert policy.
- Confirm notification or incident integrations are enabled and authorized for the destination channel, inbox, service, or webhook.
- For Slack, confirm the Bitaic app is still authorized for the destination workspace and channel.
- For Microsoft Teams, confirm the Teams Workflows webhook URL is still active and has an owner or co-owner.
- For email, confirm recipients are verified and not suppressed because of hard bounces or repeated delivery failures.
- For PagerDuty preview routing, confirm the Events API v2 routing key, service mapping, severity mapping, and
dedup_keybehavior before relying on the incident workflow. - Review whether the alert state is clear, warning, critical, resolved, or suppressed by policy.
- For webhook receivers, verify the endpoint accepts HTTPS requests and handles Bitaic delivery headers securely.
Escalation Path
Choose the highest impact level that matches the issue, then include the matching context in the support packet.
| Impact | Use when | First action |
|---|---|---|
| Critical | Production monitoring is unavailable, a critical alert route is not reaching responders, or you suspect credential, token, or data exposure. | Start your incident-response process, mitigate the affected credential or route when needed, and email support with "Critical" in the subject. |
| High | A critical target has stale telemetry, incorrect state, repeated false critical alerts, or degraded alert delivery. | Capture the latest dashboard state, run the relevant diagnostic command, and email support with "High" in the subject. |
| Normal | You need help with setup, configuration, account management, billing, product feedback, or docs feedback. | Include the product area, page URL when relevant, expected behavior, current behavior, and configuration without secrets. |
Support Escalation Packet
Include enough detail for support to reproduce the path from configuration to alert state.
| Include | Why it helps |
|---|---|
| Workspace, project, product area, and target name | Identifies the monitored object and account context. |
| Current alert state, last check time, and first observed time | Separates current issues from recovered or historical alerts. |
| Severity and customer impact | Routes active outages, alert-delivery failures, and lower-impact questions to the right response path. |
| Request ID, alert ID, delivery ID, or agent ID | Lets support correlate API responses, webhook deliveries, alerts, and agent telemetry. |
| Relevant configuration snippet with secrets removed | Lets support review thresholds, expected values, and filters. |
| CLI command output or dashboard state | Shows what Bitaic reported during the diagnostic pass. |
| Recent changes | Connects alerts to certificate renewals, DNS changes, deployments, host updates, or integration changes. |
| Expected owner and alert destination | Helps confirm whether routing or product state is the issue. |
If the issue affects production operations, people, facilities, or suspected security exposure, start your organization's incident-response process before waiting on Bitaic support. Rotate or revoke affected Bitaic API tokens, webhook secrets, and integration credentials when compromise is suspected.
Send the escalation packet to support@bitaic.com when the monitored system looks healthy but Bitaic continues to show a stale, incorrect, or unrouted state.