Skip to main content
If your webpages change on a predictable schedule, for example content that is published every working day, you can rebaseline your whole account automatically instead of doing it from the portal. This guide sets up a small Linux VM that calls the WebOrion Monitor API every weekday at 5:30 PM and rebaselines every active webpage in your account.

How it works

A short bash script makes two API calls:
  1. List webpages (GET /api/v1/urls) to get the ID of every active webpage in the account.
  2. Create batch rebaseline job (POST /api/v1/batch_jobs/rebaseline) with those IDs.
Cron runs the script on a schedule. Each run creates a batch job that you can track in the portal under Batch Jobs → View Batch Jobs, just like a bulk rebaseline started from the portal.
Rebaselining accepts the current state of each webpage as the new reference. If a webpage is defaced before the scheduled run and the alert has not been handled, the defaced version becomes the new baseline.

Before you begin

You will need:
  • An API token. Generate one in the portal under Account → Manage Credentials. The token acts as your user, so it can rebaseline any webpage your account can edit.
  • A Linux VM with outbound HTTPS (port 443) access to api.monitor.weborion.io. A small instance such as an AWS t3.micro or GCP e2-micro is enough.
The commands below are for Ubuntu 24.04. Notes for Amazon Linux 2023 are included where the steps differ.

Setup

1

Prepare the VM

Install the tools the script needs, set the timezone that the schedule should follow, and create a service user and folders for the script.
Cron uses the VM’s system timezone, so set it to the timezone you want 5:30 PM to mean. Replace Asia/Singapore if you are elsewhere.
On Amazon Linux 2023, cron is not installed by default. Use sudo dnf install -y cronie jq and then sudo systemctl enable --now crond in place of the apt command above.
2

Store the API token

Save the token in a file that only the service user can read. Keeping it out of the script means you can rotate the token without editing any code.
3

Check the token and find your webpage IDs

Before automating anything, confirm the token works by listing your webpages. Every webpage has a numeric id, which is what the rebaseline API expects.
Archived webpages are not included in this list, so they are never rebaselined by the script.To look up the ID of a single webpage by its URL, use Get a webpage by URL string:
4

Create the rebaseline script

Create /opt/weborion/rebaseline.sh with the command below. The script lists every active webpage and submits them for rebaselining in batches of 200./opt/weborion is owned by root, so the file is written with sudo tee. Keep the quotes around 'EOF' so the script is saved exactly as shown. Without them, the shell expands the script’s variables while writing the file.
Set the owner and permissions, then do a test run as the service user:
A successful run prints one line per batch job created, for example:
The test run starts a real rebaseline of every webpage in the account. Run it at a time when that is acceptable.
5

Schedule it for weekdays at 5:30 PM

Add a cron entry that runs the script at 17:30, Monday to Friday, and appends its output to a log file.
flock stops a second run from starting if the previous one is still in progress. Monday to Friday includes public holidays, so the job also runs on those days.
6

Check that it worked

After the first scheduled run, check the log on the VM:
Then check the progress of a job, either in the portal under Batch Jobs → View Batch Jobs, or with Get batch rebaseline job details. Use the numeric API ID from the log, without the RB prefix:
This returns a count of webpages per status, such as {"Completed": 198, "Error": 2}. Each webpage moves through In Queue, Processing, and then Completed or Error.

Customising the script

Rebaseline only webpages with a specific tag

If tags are enabled for your account, you can limit the run to webpages with one tag. Find the tag ID with List tags, then change the list request in the script to use List webpages by tag:

Change the schedule

Edit the first five fields of the line in /etc/cron.d/weborion-rebaseline. They are minute, hour, day of month, month and day of week.
ScheduleCron fields
Weekdays at 5:30 PM30 17 * * 1-5
Every day at 6:00 AM0 6 * * *
Mondays at 9:00 AM0 9 * * 1

Troubleshooting

SymptomLikely cause
401 from any requestThe token is wrong, expired or has been revoked. Generate a new one under Account → Manage Credentials and update /etc/weborion/rebaseline.env.
403 naming a URL IDThe user the token belongs to cannot edit that webpage.
Nothing in the log at 5:30 PMCheck that cron is running (systemctl status cron, or crond on Amazon Linux) and that the VM timezone is correct (timedatectl).
Some webpages show Error in the batch jobThose webpages could not be loaded during the rebaseline. Check that they are reachable, then rebaseline them individually from the portal.
For anything else, contact CloudsineAI support.