Docs > Platform Observability > Getting Started with Infrastructure
Getting Started with Infrastructure
Overview
This guide introduces AppStatus Infrastructure and covers:
- One-line agent install per host
- CPU, memory, disk, network and process telemetry
- Security signals: logins, privilege escalation, failed SSH
- Fleet view across every host, with per-host drill-down
- Alerting on resource pressure and on hosts that go silent
What is Infrastructure in AppStatus?
Infrastructure aggregates what the agent reports from every machine you run. The fleet view answers "is anything unhealthy right now"; each host drills down into live telemetry and running processes.
Alongside resource metrics, the agent collects security-relevant events — successful and failed logins, privilege escalation, and brute-force attempts — so a compromised host looks different from a merely busy one.
Key capabilities:
- Fleet-wide health with per-host drill-down
- CPU, memory, disk, network, uptime and process inventory
- Security signals surfaced per host
- Container activity on hosts that run them
- Silent-host detection — a host that stops reporting raises an event
What you will see
What the fleet view looks like
Scroll the table sideways to see every column.
A host that stops reporting is an event, not an absence. Treat "Not reporting" as unhealthy until proven otherwise.
Troubleshooting
The host does not appear after installing the agent
Check the agent service is running and that outbound access to AppStatus is allowed by the host firewall. The agent retries on failure, so a host that never appears is almost always blocked egress.
A host shows as not reporting but is clearly up
The agent process may have stopped or lost network access. Restart the agent, then check whether clock skew on the host is large — badly skewed clocks can make recent data look stale.
Resource alerts fire constantly on one host group
Thresholds are per group for a reason. A build machine at 90% CPU is working; a database at 90% is in trouble. Split the group and set limits that match how each is provisioned.
Operational Guidance
- A host that stops reporting is an event in itself — do not treat silence as healthy.
- Correlate a CPU spike with APM throughput before assuming the host is the problem.
- Repeated failed SSH attempts on a public host deserve a look even without an alert.
Step-by-Step Setup
One agent per host, installed with a single line. The setup that actually matters is afterwards: grouping hosts sensibly so thresholds mean something, and treating a silent host as an event.
Before you start
- Shell access to the hosts you want to monitor
- Outbound network access from those hosts
- 1
Install the agent
Copy the install command from the app and run it on the host. It installs, registers against your workspace and starts reporting — no configuration file needed for the default collectors.
WhereAgents → Setup → copy install command - 2
Confirm the host reports
Open the fleet list and confirm the host appears with live CPU, memory and disk. If it never appears, the cause is almost always blocked outbound access rather than the agent itself.
WhereInfrastructure → Fleet - 3
Group hosts by role
Group by role or environment rather than by hostname. Groups are what thresholds attach to, and a database and a build box should never share limits.
WhereInfrastructure → GroupsTipGet this right early. Regrouping later means redoing every threshold you set.
- 4
Enable security signals
Turn on security collection for anything internet-facing. It surfaces logins, privilege escalation and failed SSH attempts per host, so a compromised machine looks different from a busy one.
WhereAgents → Modules → Security - 5
Route host alerts
Send alerts to the team that owns the machine, and add an alert for hosts that stop reporting. Silence is a symptom, not an all-clear.
WhereAlerts → Alert rules
Configuration Options
Every option you can set, what each choice means, and what to pick. Use this as a reference while you fill in the form.
Agent configuration
| Field | Options | What it does | Recommended |
|---|---|---|---|
| Host group | Any grouping | Organises the fleet and carries thresholds. | Group by role or environment, never by hostname. |
| Collectors | System, security, containers | Which telemetry the agent gathers. | Enable security on anything internet-facing. |
| Thresholds | Per group | When a host counts as unhealthy. | Set against how the host is provisioned, not a global default. |
| Report interval | Seconds | How often the agent checks in. | Default suits most fleets; shorten only for volatile workloads. |
Feature Reference
Every feature, where to find it in the app, and what it does. Use this when you know what you want to do but not where it lives.
| Feature | Where in app | Description |
|---|---|---|
| Fleet view | Infrastructure → Fleet | Every host, with health at a glance and per-host drill-down. |
| Live telemetry | Host detail → Live | Current resource usage and running processes. |
| Security signals | Host detail → Security | Logins, privilege escalation and brute-force attempts. |
| Silent-host detection | Alerts | A host that stops reporting raises an event of its own. |
Next Steps
Continue building your monitoring stack:
