---
name: alibabacloud-loongcollector-ops
description: |
  Alibaba Cloud LoongCollector / SLS installation, collection onboarding, Pipeline config management and validation, machine groups, permission troubleshooting, and Lens queries.
  HARD RULE: for matching requests, the first tool MUST load this skill before any SSH probe, directory setup, checklist/file write, or cloud read.
  Triggers: "安装 LoongCollector", "ECS 安装采集器", "自建 Linux 主机安装 LoongCollector", "ACK 安装 loongcollector", "自建 K8s 部署采集", "从安装到能查到日志", "SLS 日志采集接入", "SLS 日志采集接入相关的事", "修改采集配置", "改采集配置", "采集配置校验", "validate_pipeline.py", "SLS 机器组", "新建 Logtail Pipeline 采集配置", "Logtail Pipeline", "ClusterAliyunPipelineConfig", "SLS Lens 查询", "无数据排查", "心跳异常", "SLS 采集权限排查", "SLS 权限排查", "阿里云 CLI 凭证没有 SLS 操作权限", "Logtail", "iLogtail", "AgentSight", "Agentloop", "input_agentsight", "eBPF Runtime", "ebpf-event".
---

# LoongCollector Ops

Turn natural-language requests into executable, verifiable, rollbackable workflows for users operating **their own** LoongCollector and SLS resources — from install through collection to query.

**Architecture**: `Install (ECS/self-host/ACK/self-k8s) + SLS Project + Logstore + Index + MachineGroup + Pipeline (API or ClusterAliyunPipelineConfig) + binding + SLS Lens`

**Scope.** Covers:
- Install/upgrade Linux collector: ECS `aliyun ecs run-command`, self-host SSH, ACK addon `loongcollector`, self-k8s custom package. Must continue to collection + query; process/Addon ready is only a stage gate.
- ACK first-use: `open-ack-service --type propayasgo` + CS service roles (`scripts/ensure_ack_prereq.sh`). `create-cluster` only when the user asked to create a cluster. Eval hooks may pre-create a fixture cluster; that is not a production default.
- Cloud onboarding: Project / Logstore / Index / MachineGroup / Pipeline Config / binding.
- K8s collection: **default SLS Pipeline API** (`create-logtail-pipeline-config` + bind official group). CRD apply is opt-in only when the user asks for GitOps/CRD and a reachable kube-apiserver exists (`references/crd-pipeline.md`).
- Config management: create, modify, apply, remove, data acceptance (U1-U6).
- Machine group management: IP / user-defined identity, members, heartbeat, version.
- SLS Lens: run-log query (`get-logs-v2`), topic/field contracts, version routing, degradation.
- Basic troubleshooting: no-data, heartbeat abnormal.

## Language and HITL Delivery Contract

**Hard language rule:** use the user's primary language for every user-facing message. This includes plans, clarification questions, confirmation questions and their answer options, reports, and error guidance. Product names, identifiers, CLI commands, JSON fields, error codes, and fixed status tags may remain in their original form. Never switch the surrounding prose to another language.

**Canonical user-facing message and marker catalog:** every value below is literal. Emit the selected value verbatim; never translate, paraphrase, or combine it with another question.

```yaml
messages:
  missing_task_scope: "请补充要执行的具体操作目标、地域和 SLS Project。"
  missing_lens_parameters: "请补充业务 Project、地域和查询时间范围。"
  machine_group_identity: "请选择机器组标识类型：IP 或 userdefined。"
  r2_update: "是否确认执行上述变更计划？请选择：确认执行或取消。"
  r2_create_bind: "是否确认创建上述资源并完成绑定？请选择：确认执行或取消。"
  r3_unbind: "是否确认将上述旧配置从机器组解绑？请选择：确认解绑或取消。"
  permission_recovery: "是否已完成所需 RAM 授权并允许重试？请选择：已授权或未授权。"
  permission_recovery_short: "是否已完成所需 RAM 授权并允许重试？"
  lens_entry: "请提供 SLS Lens 服务日志的 Project 和 Logstore。"
  ecs_install: "是否确认在上述 ECS 上安装 LoongCollector？"
  self_host_install: "是否确认在上述主机上安装 LoongCollector？"
  ack_install: "是否确认在上述 ACK 集群安装 loongcollector 组件？"
  self_k8s_install: "是否确认在上述 Kubernetes 集群安装 LoongCollector？"
  kubeconfig: "请提供可用的 kubectl 与目标集群 context。"
  ssh: "请提供已配置的 SSH（alias 或主机），不要在对话中发送私钥。"
  collector_version: "请提供采集器版本（例如 3.3.9）。"
markers:
  ownership_error: ["不属于当前账号", "项目不属于你"]
  permission_decline: ["未授权", "停止"]
  permission_grant: ["已授权"]
  cancel: ["取消"]
  private_ip: ["私网 IP"]
  routing_intent: ["日志采集", "安装采集器"]
  end: ["结束"]
  collector_deployed_without_version: ["LoongCollector 已部署"]
  binding_acceptance: ["完成绑定与验收"]
  approval: ["确认", "确认执行", "确认解绑"]
  install_intent: ["允许安装", "请安装", "直接执行", "已授权操作", "任务已预授权"]
  install_only_status: ["仅安装完成、采集未接入"]
  deferral: ["还没想好", "等会儿再说", "暂不确认", "第二次等待", "第N次暂不确认", "先放一放", "已达到上限", "请阻塞"]
  data_incomplete: ["无法完成数据面验收"]
  data_empty: ["无数据"]
  reason: ["原因"]
  forbidden_empty_success: ["采集成功", "所有验收标准均已满足", "全链路验收通过", "通过"]
  data_arrived: ["数据到达"]
  root_cause_located: ["根因已定位"]
  not_exists: ["不存在"]
  pending_read: ["未执行待办", "待办"]
```

Pair `machine_group_identity` with `[AWAITING: MACHINE_GROUP_TYPE]` (never `R2_CONFIRMATION`). Pair `permission_recovery` with `[AWAITING: PERMISSION_CONFIRMATION]`; `lens_entry` with `[AWAITING: LENS_ENTRY]`; every install message with `[AWAITING: INSTALL_CONFIRMATION]`; `kubeconfig` with `[AWAITING: KUBECONFIG]`; `ssh` with `[AWAITING: SSH]`; and `collector_version` with `[AWAITING: COLLECTOR_VERSION]`. The `self_host_install` message is allowed only after a real SSH probe succeeds.

Do not replace these with long English prose, bilingual tables, or newly invented status labels. Whenever you re-ask, reproduce the same short Chinese question verbatim before the required `[AWAITING: ...]` tag. **Last-line hard rule:** the matching tag immediately follows the question on the next line and is the last line of the turn — no blank line between question and tag, no blank line after it, no punctuation, and no extra sentence. The turn that emits a HITL tag must not copy any `[AWAITING: ...]` literal into a tool call, code block, `outputs/*`, or `ran_scripts/*`; duplicate tags break automatic matching. Install confirmation ends the turn with `[AWAITING: INSTALL_CONFIRMATION]` — never reuse `R2_CONFIRMATION` for install or for machine-group identity. Collection/create-bind confirmation tags MUST include the ask counter on the last line: first ask `[AWAITING: R2_CONFIRMATION] ask=1`; each deferral re-ask increments the counter. Lens-entry fallback ends the turn with `[AWAITING: LENS_ENTRY]`. Missing kubectl ends the turn with `[AWAITING: KUBECONFIG]`. Missing SSH ends the turn with `[AWAITING: SSH]`. Missing collector version ends the turn with `[AWAITING: COLLECTOR_VERSION]`. Machine-group identity ends with `[AWAITING: MACHINE_GROUP_TYPE]`. RAM recovery ends with `[AWAITING: PERMISSION_CONFIRMATION]`.

**Fixed English tokens (must appear verbatim; surrounding prose stays Chinese):** `[BLOCKED: …]` / `[CANCELLED: …]` / `[AWAITING: …]` / `ask=1` / `ask=2` / `ask=3` / `[Error: permission|throttling|internal|parameter]` / `[RECOVERED: …]` / `resource_status: Resource not found` / `[Query: Incomplete]` / `INCOMPLETE`. Rejection and confirmation-timeout turns: the **sole content** of that turn is the short tag — no English long sentence, no prefix or suffix.

**Out of scope.** Windows; Sidecar; uninstall/rollback/restart-as-lifecycle; creating ECS; OOS/ChatOps; writing `AliyunLogConfig` / `NamespaceAliyunPipelineConfig`; advanced troubleshooting (delay, duplicate, parse failure, container filter, data loss/truncation). `kubectl exec` and `docker exec` are forbidden. If the user asks for an out-of-scope lifecycle action, say so and stop that branch.

---

## 1. Prerequisites

**Pre-check: Aliyun CLI >= 3.3.3 required**
> [MUST] Verify: `aliyun version` — must be >= 3.3.3 (>= 3.3.5 recommended).
> - First install or major upgrade: `/bin/bash -c "$(curl -fsSL --connect-timeout 10 --max-time 120 https://aliyuncli.alicdn.com/setup.sh)"`
> - Routine update (CLI >= 3.3.5): `aliyun upgrade`.
> - See `references/cli-installation-guide.md`.

**Pre-check: SLS plugin required**
> [MUST] `aliyun configure set --auto-plugin-install true` then `aliyun plugin install --names aliyun-cli-sls` and `aliyun plugin update`.
> Collection subcommands are provided by the `aliyun-cli-sls` plugin (hyphenated subcommands such as `aliyun sls get-logs-v2`). Verify with `aliyun sls --help`.

> **Pre-check: Alibaba Cloud Credentials Required**
>
> **Security Rules:**
> - **NEVER** read, echo, or print AK/SK values (e.g., `echo $ALIBABA_CLOUD_ACCESS_KEY_ID` is FORBIDDEN)
> - **NEVER** use `cat`, `less`, `head`, `tail`, `grep`, `open`, `json.load`, or any file-reading command on credential files (e.g., `~/.aliyun/config.json`, `~/.aws/credentials`). To check file existence use `ls` only — never display contents. Printing plaintext secrets is an immediate task failure and security incident.
> - **NEVER** install or import `aliyun-log-python-sdk` / `aliyun.log` / `LogClient`, or any other SLS SDK, to bypass CLI. `pip install` of a cloud SDK is a task failure.
> - **NEVER** print, `cat`, or paste kubeconfig / client certificates / tokens into the conversation. For opt-in CRD only: write `describe-cluster-user-kubeconfig` output to a `0600` tempfile.
> - **NEVER** ask the user to input AK/SK directly in the conversation or command line
> - **NEVER** use `aliyun configure set` with literal credential values
> - **ONLY** use `aliyun configure list` to check credential status. `scripts/preflight.sh` already does this.
>
> ```bash
> aliyun configure list
> ```
> Check the output for a valid profile (AK, STS, or OAuth identity).
>
> **If no valid profile exists, STOP here.**
> 1. Obtain credentials from [Alibaba Cloud Console](https://ram.console.aliyun.com/manage/ak)
> 2. Configure credentials **outside of this session** (via `aliyun configure` in terminal or environment variables in shell profile)
> 3. Return and re-run after `aliyun configure list` shows a valid profile

Run `bash scripts/preflight.sh` to check CLI version, plugin, credential presence, and scope in one step. `preflight.sh` already invokes `aliyun configure list` internally; running it **is** a valid credential check — do not cat CLI config files, and do not add a standalone `aliyun configure list` just to satisfy a checklist. Full gate details: `references/prerequisites.md`.

### Environment Variables

| Variable | Required | Description |
|---|---|---|
| (none for credentials) | — | Credentials come from `aliyun configure` profiles; never introduce AK/SK env vars in-session |
| `SKILL_SESSION_ID` | Injected at script run | Same 32-hex session id as the `session/{session-id}` UserAgent token; set inline when invoking bundled scripts (see §4) |

---

## 2. RAM Policy

This skill uses the user's own identity and only touches resources they are authorized for. Permissions are layered ReadOnly / Operator / Destructive. Per-workflow RAM Actions are in `references/ram-policies.md` — do not default to broad `AliyunLogFullAccess`.

> **[MUST] Permission Failure Handling:** When any command or API call fails due to permission errors at any point during execution, follow this process:
> 1. Read `references/ram-policies.md` to get the full list of permissions required by this SKILL
> 2. Use `ram-permission-diagnose` skill to guide the user through requesting the necessary permissions
> 3. Pause and wait until the user confirms that the required permissions have been granted

**Runtime detail (same gate, do not skip the three steps above):**
1. Report the missing RAM Action and `requestID`; output `[Error: permission]`. Reading `references/ram-policies.md` alone is **not** a successful diagnose call.
2. **Try** `ram-permission-diagnose` with the missing Actions and `requestID`. **FALLBACK:** if it is unavailable, output Action/`requestID`/RAM-console guide manually, then ask catalog message `permission_recovery` with last line `[AWAITING: PERMISSION_CONFIRMATION]` and pause.
3. Do not retry the affected write (including `--cli-dry-run`) before confirmation.
4. **READ-PATH HARD STOP:** On 401/403/`Unauthorized`/`AccessDenied` for `get-project` / `get-machine-group` / `list-machines` / `get-log-store`, first read the message. If it is **ownership** (English ownership text or any catalog `ownership_error` marker) → this is **not** a RAM gate: emit `[BLOCKED: RESOURCE_RESOLUTION_FAILED]` and stop; do **not** ask `permission_recovery_short`, do not create the official `k8s-log-*` name, and do not retry. Otherwise emit `[Error: permission]` with Action/`requestID`, then in the **same turn** ask exactly `permission_recovery` with last line `[AWAITING: PERMISSION_CONFIRMATION]`, and issue **zero** further `aliyun sls` calls that turn — including `get-machine-group`, `list-machines`, `get-applied-configs`, and `get-log-store`. Those unread calls are catalog `pending_read` items, not queried conclusions.
5. **After the user's permission answer (same gate for read-path and write/dry-run):**
   - Any catalog `permission_decline` marker or equivalent decline → **zero tools that turn** (no `write_file`, no `aliyun sls`). Explicitly state that execution is terminated, list every **not-yet-run** read as catalog `pending_read` with its RAM Action, then put `[BLOCKED: PERMISSION_REQUIRED]` on the final line. If the task asked for machine-group heartbeat, the pending items **must** include `get-machine-group` → `log:GetMachineGroup` and `list-machines` → `log:ListMachines` (name the group). Never present an unrun heartbeat as a queried conclusion.
   - A catalog `permission_grant` marker → **same turn**, retry the **identical** failed command (if the failure was a dry-run, retry that dry-run first) and emit `[RECOVERED: permission_granted]` in the user-facing text immediately. Explicitly state that the disposition is human intervention followed by retry, so the recovery action is unambiguous.

On `Unauthorized`/`AccessDenied` from a **core write or its dry-run**: stop the current write, enter the §6 permission-recovery branch, and never switch account/profile or widen scope.

---

## 3. Parameter Confirmation

> **IMPORTANT: Parameter Confirmation** — Before executing any command or API call,
> ALL user-customizable parameters (e.g., RegionId, instance names, CIDR blocks,
> passwords, domain names, resource specifications, etc.) MUST be confirmed with the
> user. Do NOT assume or use default values without explicit user approval.

| Parameter | Required/Optional | Description | Default |
|---|---|---|---|
| `region` | Required | SLS region id, e.g. `cn-hangzhou` | none |
| `project` | Required | Target SLS project | none |
| `logstore` | Conditional | Target logstore (create/onboarding/query) | none |
| `machine_group` | Conditional | Machine group name (bind/heartbeat) | none |
| `config_name` | Conditional | Logtail pipeline config name | none |
| `machine_identify_type` | Conditional | `ip` or `userdefined` | none |
| `machine_list` | Conditional | IP list or user-defined identifiers | none |
| `scenario` | Conditional | `host` / `docker` / `k8s` / `host_agentsight` | none |
| `environment` | Conditional | `ecs` / `self_host` / `ack` / `self_k8s` for `install.deploy` | none |
| `instance_id` | Conditional | ECS instance id (Workbench) | none |
| `cluster_id` | Conditional | ACK / self-k8s cluster id | none |
| `lens_project` / `lens_logstore` | Optional | SLS Lens entry when auto-discovery unavailable | none |

Never substitute placeholder/example values for a missing real resource. Missing `region`, `project`, or a scope-changing target → STOP and ask.

For a Chinese request with a missing `machine_identify_type`, ask exactly catalog message `machine_group_identity` using Chinese option descriptions, and end the turn with last line `[AWAITING: MACHINE_GROUP_TYPE]` — **never** `[AWAITING: R2_CONFIRMATION]`. After the user chooses, run `scripts/normalize_diff.py` and only then ask the create-and-bind question with `[AWAITING: R2_CONFIRMATION] ask=1`. Do not emit a custom English status label or an English selection table. **`install.deploy` exception:** do **not** ask machine-group identity until the collection create-bind gate has been confirmed. If that gate receives a catalog `cancel` marker or equivalent rejection, never ask. If the user already named IP, userdefined, or a catalog `private_ip` marker, never ask.

---

## 4. Observability (MUST follow for every aliyun command)

Upon loading this skill, generate a random session ID (32-char lowercase hex string) once for the entire session. Use it as `{session-id}` below.

**Rule: Every `aliyun` CLI command that calls a cloud API MUST include the `--user-agent` flag.**
This covers `aliyun sls`, `aliyun cs`, `aliyun ecs`, and any other `aliyun <product>` cloud call this skill issues, including `--cli-dry-run`. Local utility commands (e.g. `configure`, `plugin`, `version`) do not support this flag and should be excluded. `kubectl` / Workbench / SSH / local validators are not Alibaba Cloud APIs and do not send this flag.

Use **two space-separated product tokens** (quote the whole value; the space is required):

```
--user-agent "AlibabaCloud-Agent-Skills/alibabacloud-loongcollector-ops session/{session-id}"
```

| Token | Example | Query use |
|---|---|---|
| Skill identity | `AlibabaCloud-Agent-Skills/alibabacloud-loongcollector-ops` | All traffic from this skill |
| Session | `session/{session-id}` | One session |

Never glue the session id onto the skill token (`.../ops/{session-id}` is forbidden). Never omit quotes. Never skip, alter, or drop either token.

Example (assuming session-id is `a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6`):
```bash
aliyun sls list-machines --project my-proj --machine-group my-group --region cn-hangzhou --user-agent "AlibabaCloud-Agent-Skills/alibabacloud-loongcollector-ops session/a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6"
```

References that write `--user-agent <ua>` mean this exact quoted two-token string.

**Script / Terraform execution:** When running Python SDK scripts or Terraform commands or bash scripts, inject the session-id via inline environment variable so the code can read it at runtime:

```bash
# Local validator (no cloud call)
SKILL_SESSION_ID={session-id} python3 scripts/validate_pipeline.py --file rendered.json

# Bundled script that itself calls a cloud API
SKILL_SESSION_ID={session-id} bash scripts/wait_cs_task.sh --cluster-id c-xxx --region cn-shanghai

# Terraform
SKILL_SESSION_ID={session-id} terraform apply
```

Scripts and Terraform configs should read `SKILL_SESSION_ID` from the environment (default to empty string if absent). Any bundled script that itself invokes `aliyun` against a cloud API MUST send the same two-token UserAgent (currently `scripts/wait_cs_task.sh`).

**Domain extension — ATOMIC CLOUD-CALL RULE (HARD):** Every tool invocation that calls SLS must contain exactly one direct `aliyun sls ...` command with literal, fully expanded parameter values, and the command must start with `aliyun sls`. Do not hide a cloud call behind shell variables, environment assignments, functions, aliases, wrapper scripts, loops, command substitutions, `eval`, pipes, redirections (including `2>&1`), or compound commands (`;`, `&&`, `||`). Use only lowercase hyphenated SLS plugin subcommands such as `get-project`; never use a PascalCase OpenAPI alias, because it bypasses the validated command/mocking contract. Compute timestamps or JSON in a separate local step, then place the resulting literals in the cloud command. This applies equally to verification and acceptance reads: no loops, no `cd …` prefix, no `$VAR` or `$(…)` substitution, including in `--from`/`--to` and JSON bodies. Never write cloud calls into a `.sh` file and run it; the command record is plain-text notes, not a runnable script. Local validators must receive what the command actually returned — save the real stdout to a file and pass that file; retyping or `echo`-ing an expected response is fabricated evidence. Generate the session ID with `python3 -c 'import secrets; print(secrets.token_hex(16))'` and validate `^[0-9a-f]{32}$`; never copy the example ID above into live commands.

---

## 5. Capability Router

Classify the request into exactly one capability, then load its `references/navigation.md` entry before acting. Do not load the whole knowledge base in