> ## Documentation Index
> Fetch the complete documentation index at: https://docs.weborion.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Scheduled Account Rebaseline

> Use the WebOrion Monitor API to rebaseline every webpage in your account on a fixed schedule.

If your webpages change on a predictable schedule, for example content that is published every working day, you can rebaseline your whole account automatically instead of doing it from the portal.

This guide sets up a small Linux VM that calls the WebOrion Monitor API every weekday at 5:30 PM and rebaselines every active webpage in your account.

## How it works

A short bash script makes two API calls:

1. [List webpages](/api-reference/webpages/list-webpages) (`GET /api/v1/urls`) to get the ID of every active webpage in the account.
2. [Create batch rebaseline job](/api-reference/batch-jobs/create-batch-rebaseline-job) (`POST /api/v1/batch_jobs/rebaseline`) with those IDs.

Cron runs the script on a schedule. Each run creates a batch job that you can track in the portal under **Batch Jobs → View Batch Jobs**, just like a bulk rebaseline started from the portal.

<Warning>
  Rebaselining accepts the current state of each webpage as the new reference. If a webpage is defaced before the scheduled run and the alert has not been handled, the defaced version becomes the new baseline.
</Warning>

## Before you begin

You will need:

* **An API token.** Generate one in the portal under **Account → Manage Credentials**. The token acts as your user, so it can rebaseline any webpage your account can edit.
* **A Linux VM** with outbound HTTPS (port 443) access to `api.monitor.weborion.io`. A small instance such as an AWS `t3.micro` or GCP `e2-micro` is enough.

The commands below are for Ubuntu 24.04. Notes for Amazon Linux 2023 are included where the steps differ.

## Setup

<Steps>
  <Step title="Prepare the VM">
    Install the tools the script needs, set the timezone that the schedule should follow, and create a service user and folders for the script.

    ```bash theme={null}
    sudo apt update && sudo apt install -y curl jq cron
    sudo timedatectl set-timezone Asia/Singapore
    sudo useradd --system --create-home --shell /usr/sbin/nologin weborion
    sudo mkdir -p /opt/weborion /etc/weborion /var/log/weborion
    sudo chown weborion:weborion /var/log/weborion
    ```

    Cron uses the VM's system timezone, so set it to the timezone you want 5:30 PM to mean. Replace `Asia/Singapore` if you are elsewhere.

    <Note>
      On Amazon Linux 2023, cron is not installed by default. Use `sudo dnf install -y cronie jq` and then `sudo systemctl enable --now crond` in place of the `apt` command above.
    </Note>
  </Step>

  <Step title="Store the API token">
    Save the token in a file that only the service user can read. Keeping it out of the script means you can rotate the token without editing any code.

    ```bash theme={null}
    sudo tee /etc/weborion/rebaseline.env >/dev/null <<'EOF'
    WEBORION_API_TOKEN=paste-your-token-here
    EOF
    sudo chown weborion:weborion /etc/weborion/rebaseline.env
    sudo chmod 600 /etc/weborion/rebaseline.env
    ```
  </Step>

  <Step title="Check the token and find your webpage IDs">
    Before automating anything, confirm the token works by listing your webpages. Every webpage has a numeric `id`, which is what the rebaseline API expects.

    ```bash theme={null}
    export WEBORION_API_TOKEN=paste-your-token-here

    curl -sS -H "Authorization: Bearer $WEBORION_API_TOKEN" \
      https://api.monitor.weborion.io/api/v1/urls \
      | jq '(if type=="array" then . else .data end)[] | {id, url_name, config_name, health_status}'
    ```

    Archived webpages are not included in this list, so they are never rebaselined by the script.

    To look up the ID of a single webpage by its URL, use [Get a webpage by URL string](/api-reference/webpages/get-a-webpage-by-url-string):

    ```bash theme={null}
    curl -sS -X POST -H "Authorization: Bearer $WEBORION_API_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{"url": "https://www.example.com/"}' \
      https://api.monitor.weborion.io/api/v1/urls/get_by_url | jq '.id'
    ```
  </Step>

  <Step title="Create the rebaseline script">
    Create `/opt/weborion/rebaseline.sh` with the command below. The script lists every active webpage and submits them for rebaselining in batches of 200.

    `/opt/weborion` is owned by root, so the file is written with `sudo tee`. Keep the quotes around `'EOF'` so the script is saved exactly as shown. Without them, the shell expands the script's variables while writing the file.

    ```bash theme={null}
    sudo tee /opt/weborion/rebaseline.sh >/dev/null <<'EOF'
    #!/usr/bin/env bash
    # Rebaseline every active webpage in a WebOrion Monitor account.
    set -euo pipefail

    source "${ENV_FILE:-/etc/weborion/rebaseline.env}"
    : "${WEBORION_API_TOKEN:?WEBORION_API_TOKEN is not set}"
    API_BASE="${WEBORION_API_BASE:-https://api.monitor.weborion.io}"
    BATCH_SIZE="${BATCH_SIZE:-200}"

    log() { echo "$(date '+%Y-%m-%d %H:%M:%S %Z') $*"; }

    api() {
      curl -sS --fail-with-body --retry 3 --retry-delay 5 --max-time 120 \
        -H "Authorization: Bearer ${WEBORION_API_TOKEN}" \
        -H "Accept: application/json" \
        -H "Content-Type: application/json" \
        "$@"
    }

    log "Fetching webpages"
    urls_json="$(api "${API_BASE}/api/v1/urls")"

    total="$(jq '(if type=="array" then . else .data end) | length' <<<"$urls_json")"
    if [[ "$total" -eq 0 ]]; then
      log "No active webpages found, nothing to do"
      exit 0
    fi
    log "Found ${total} active webpages"

    jq -c --argjson n "$BATCH_SIZE" \
      '[(if type=="array" then . else .data end)[].id]
       | [range(0; length; $n) as $i | .[$i:$i+$n]][]' <<<"$urls_json" |
    while read -r chunk; do
      count="$(jq length <<<"$chunk")"
      resp="$(api -X POST "${API_BASE}/api/v1/batch_jobs/rebaseline" -d "{\"url_ids\": ${chunk}}")"
      job_id="$(jq -r '.id' <<<"$resp")"
      log "Created rebaseline job RB$(printf '%05d' "$job_id") (API ID ${job_id}) for ${count} webpages"
    done

    log "Done"
    EOF
    ```

    Set the owner and permissions, then do a test run as the service user:

    ```bash theme={null}
    sudo chown weborion:weborion /opt/weborion/rebaseline.sh
    sudo chmod 750 /opt/weborion/rebaseline.sh
    sudo -u weborion /opt/weborion/rebaseline.sh
    ```

    A successful run prints one line per batch job created, for example:

    ```text theme={null}
    2026-10-01 17:30:01 +08 Fetching webpages
    2026-10-01 17:30:02 +08 Found 312 active webpages
    2026-10-01 17:30:03 +08 Created rebaseline job RB01234 (API ID 1234) for 200 webpages
    2026-10-01 17:30:03 +08 Created rebaseline job RB01235 (API ID 1235) for 112 webpages
    2026-10-01 17:30:03 +08 Done
    ```

    <Note>
      The test run starts a real rebaseline of every webpage in the account. Run it at a time when that is acceptable.
    </Note>
  </Step>

  <Step title="Schedule it for weekdays at 5:30 PM">
    Add a cron entry that runs the script at 17:30, Monday to Friday, and appends its output to a log file.

    ```bash theme={null}
    sudo tee /etc/cron.d/weborion-rebaseline >/dev/null <<'EOF'
    30 17 * * 1-5 weborion flock -n /tmp/weborion-rebaseline.lock /opt/weborion/rebaseline.sh >> /var/log/weborion/rebaseline.log 2>&1
    EOF
    sudo chmod 644 /etc/cron.d/weborion-rebaseline
    ```

    `flock` stops a second run from starting if the previous one is still in progress. Monday to Friday includes public holidays, so the job also runs on those days.
  </Step>

  <Step title="Check that it worked">
    After the first scheduled run, check the log on the VM:

    ```bash theme={null}
    tail -n 20 /var/log/weborion/rebaseline.log
    ```

    Then check the progress of a job, either in the portal under **Batch Jobs → View Batch Jobs**, or with [Get batch rebaseline job details](/api-reference/batch-jobs/get-batch-rebaseline-job-details). Use the numeric API ID from the log, without the `RB` prefix:

    ```bash theme={null}
    curl -sS -H "Authorization: Bearer $WEBORION_API_TOKEN" \
      https://api.monitor.weborion.io/api/v1/batch_jobs/rebaseline/1234 \
      | jq '[.entries[].status] | group_by(.) | map({(.[0]): length}) | add'
    ```

    This returns a count of webpages per status, such as `{"Completed": 198, "Error": 2}`. Each webpage moves through `In Queue`, `Processing`, and then `Completed` or `Error`.
  </Step>
</Steps>

## Customising the script

### Rebaseline only webpages with a specific tag

If tags are enabled for your account, you can limit the run to webpages with one tag. Find the tag ID with [List tags](/api-reference/tags/list-tags), then change the list request in the script to use [List webpages by tag](/api-reference/tags/list-webpages-by-tag):

```bash theme={null}
urls_json="$(api "${API_BASE}/api/v1/tags/<tag_id>/urls")"
```

### Change the schedule

Edit the first five fields of the line in `/etc/cron.d/weborion-rebaseline`. They are minute, hour, day of month, month and day of week.

<table>
  <colgroup>
    <col width="197" />

    <col width="521" />
  </colgroup>

  <thead>
    <tr>
      <th>Schedule</th>
      <th>Cron fields</th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td>Weekdays at 5:30 PM</td>
      <td>`30 17 * * 1-5`</td>
    </tr>

    <tr>
      <td>Every day at 6:00 AM</td>
      <td>`0 6 * * *`</td>
    </tr>

    <tr>
      <td>Mondays at 9:00 AM</td>
      <td>`0 9 * * 1`</td>
    </tr>
  </tbody>
</table>

## Troubleshooting

<table>
  <colgroup>
    <col width="256" />

    <col width="465" />
  </colgroup>

  <thead>
    <tr>
      <th>Symptom</th>
      <th>Likely cause</th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td>`401` from any request</td>
      <td>The token is wrong, expired or has been revoked. Generate a new one under **Account → Manage Credentials** and update `/etc/weborion/rebaseline.env`.</td>
    </tr>

    <tr>
      <td>`403` naming a URL ID</td>
      <td>The user the token belongs to cannot edit that webpage.</td>
    </tr>

    <tr>
      <td>Nothing in the log at 5:30 PM</td>
      <td>Check that cron is running (`systemctl status cron`, or `crond` on Amazon Linux) and that the VM timezone is correct (`timedatectl`).</td>
    </tr>

    <tr>
      <td>Some webpages show `Error` in the batch job</td>
      <td>Those webpages could not be loaded during the rebaseline. Check that they are reachable, then rebaseline them individually from the portal.</td>
    </tr>
  </tbody>
</table>

For anything else, contact [CloudsineAI support](/defacement-monitor/getting-support/customer-support).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.