Skip to content

Free tool

alarm-recipient-scan — what happens when the far end is down

Anyone who turns Niagara alarms into something outside Niagara — a work order, a ticket, a message, a webhook — implements one method on BRecoverableRecipient, and two words in its signature decide what happens on a bad day. alarm-recipient-scan reads alarm-rt.jar and baja.jar with javap and prints what the retry machinery really does, what reaches the log, and which properties hold the evidence.

  • Niagara alarms
  • Recipients
  • Retry and recovery
  • No station needed
  • Read-only
  • MIT

Install and run

One file, standard library only, and nothing on the network: it reads two jars on disk. It needs a Niagara installation and a javap on the path, which the JDK that ships with Niagara already provides. Point it at the installation or set NIAGARA_HOME.

# download it next to wherever you are working
curl -O https://plantroomlabs.com/tools/alarm-recipient-scan.py

# read the recipient machinery out of an installation
python3 alarm-recipient-scan.py /opt/Niagara/Niagara-4.15.5.22

What it printed here

This is the whole of one run against the installation on the machine that built this page, pasted by the script that publishes it rather than retyped. The method every custom recipient implements is protected abstract boolean sendAlarm(BAlarmRecord) throws Exception, and the output is mostly about those two words.

Returning false loses the alarm. It reads like “not sent, please retry”. In the bytecode it is a bare return: the record is not queued, the status is not set to fault, lastFailureCause is not written, the retry thread is not started, and nothing is logged. The alarm is gone and the station still looks healthy. Throwing is the only way into the recovery machinery — the exception handler is what sets fault status, persists the record and starts the retry loop.

There is no give-up and no dead letter. A record the far end will never accept is retried on a 15-second default for as long as the station runs, and with persistent=true each retry is one file delete and one file write on a controller's flash.

Almost none of it reaches the log. Of the seven logging sites in the class, six are FINE — off by default — and only the persistence failure is a WARNING. The evidence a site can actually send you lives in the recipient's own properties: status, lastFailureTime, lastFailureCause and queuedAlarmCount. A product that does not surface those four leaves an operator with nothing to look at.

$ python3 alarm-recipient-scan.py $NIAGARA_HOME
Alarm recipients: what happens when the far end is down
======================================================================

read from /opt/Niagara/Niagara-4.15.5.22
javap:    1.8.0_504

The one method a vendor implements:
  protected abstract boolean sendAlarm(javax.baja.alarm.BAlarmRecord) throws java.lang.Exception;

Property defaults, read out of the static initialiser, not the docs:
  retryInterval  15000 ms (15 s), with a 'min' facet of 1000 ms
  persistent     true
  so out of the box a failed alarm is written to disk and retried
  every 15 seconds, for as long as the station runs.

Step 1. handleAlarm calls sendAlarm on the alarm thread, at offset 82.
        The return value is tested once:
          ifne 91   -> true: set lastSendTime, status ok
          90: return  <- false: RETURN, and nothing else at all

  So returning false does not mean 'not sent, please retry'. It means
  'forget this alarm'. The record is not queued, the status is not set
  to fault, lastFailureCause is not written, the retry thread is not
  started, and nothing is logged. The alarm is gone, and the station
  looks healthy.

Step 2. A thrown Exception is the only way into the recovery machinery.
        The handler (ranges 40-90, 91-145) does, in order:
          status := fault, lastFailureTime := now,
          lastFailureCause := the exception's toString (its cause, if
          it is a BajaRuntimeException),
          then if persistent: write <uuid>.xml under the recipient's
          persistence directory, inside AccessController.doPrivileged,
          and set queuedAlarmCount from a listing of that directory;
          else: Queue.enqueue the record in memory,
          then start the RetryThread if it is not already running.

Step 3. The in-memory queue is the no-argument javax.baja.util.Queue
        (the recipient calls "<init>":()V), whose default bound is
        2,147,483,647 - Integer.MAX_VALUE.
        public synchronized boolean enqueue(java.lang.Object) throws javax.baja.util.QueueFullException;
        throws only at that bound, and handleAlarm does not catch it.
        On a JACE the real bound is the heap, not a setting.

Step 4. The retry loop, in a thread named 'alarm:RecipRetryThread':
          sleep(Math.max(retryInterval, 1000 ms)) - so a retryInterval
          under 1 s is clamped up to 1 s - then poll().
          poll(): persistent and queuedAlarmCount > 0 -> dequeueDisk,
          otherwise dequeueMemory. When queuedAlarmCount reaches 0 the
          thread kills itself and the field is nulled, so the loop
          exists exactly as long as there is something to retry.
          dequeueMemory takes a size snapshot, dequeues that many and
          calls handleAlarm on each, so a record that fails again is
          re-queued at the tail: the retry order is not the alarm order.
          dequeueDisk lists the directory, sorts it with a comparator,
          and for each file decodes the record, deletes the file, then
          calls handleAlarm - which re-writes the same file on failure.

Step 5. What reaches the station log: 6 FINE sites and 1 WARNING.
        The messages in the class are:
          RecoverableRecipient: Failed to delete queue file
          RecoverableRecipient: failed @
          RecoverableRecipient: failed to persist alarm
          RecoverableRecipient: handleAlarm
          RecoverableRecipient: sending ...
          RecoverableRecipient: sent @
        Only the persistence failure is a WARNING. Every send failure,
        every retry and every poll error is FINE, which is off by
        default - so the evidence lives in the recipient's own
        properties (status, lastFailureTime, lastFailureCause,
        queuedAlarmCount), not in the log a site will send you.

Three consequences worth designing for
----------------------------------------------------------------------
  1. return false only for 'this alarm is not mine'. For a failure,
     throw. Returning false on a timeout loses alarms silently.
  2. There is no give-up and no dead letter. A record the far end will
     never accept is retried every 15 s for as long as the station
     runs. With persistent=true that is one delete and one write per
     retry - about 11520 a day - on a controller's flash.
  3. queuedAlarmCount and lastFailureCause are the only operator-
     visible evidence. A product that does not surface them leaves a
     site with nothing to look at.

Three checks one station settles in an afternoon
----------------------------------------------------------------------
  1. Point the recipient at a dead endpoint, raise one alarm, and read
     queuedAlarmCount: zero means your sendAlarm returned false and the
     alarm is gone; one means it threw and the alarm is queued.
  2. Leave it failing for ten minutes and list the persistence
     directory: one <uuid>.xml per queued alarm, rewritten every 15 s.
  3. Bring the endpoint back and watch the order the records arrive in.

What it reads, and what that does not cover

It reads alarm-rt.jar and baja.jar with javap and nothing else. No station is contacted, no alarm is sent and nothing is written. It is a reading of compiled code, so it says what the framework is built to do rather than what your recipient did at 03:00. The checks at the end of the output are the ones one station and one deliberately broken endpoint settle in an afternoon.

The numbers above came from one installation, 4.15.5.22, and the output names it. The controller this work usually targets runs 4.14.0.162. A different Niagara version is a different answer, so run it against yours rather than trusting this page.

alarm-route-scan reads the other end of the same path — the queue and the coalesce rule that decide whether a recipient is ever asked — and why an alarm never reached anybody is the configuration side of it.

The repository

The same file, MIT licensed, at github.com/UsamaIqbal0304/alarm-recipient-scan. Its README carries the findings and a run of the output, so the repository is checkable without downloading anything. Issues and pull requests are read.

The file you are downloading

Published here so the download is checkable rather than trusted. Both figures are read off the file served at /tools/alarm-recipient-scan.py when this page is built, so they cannot disagree with it.

PropertyValue
Filealarm-recipient-scan.py
Size13,607 bytes
SHA-256 e1c6741248896d16ccbf567bac5cc7004cbbb831956e7e773a57cead4ccbb23a
Licence MIT — LICENSE.txt
Source github.com/UsamaIqbal0304/alarm-recipient-scan

To check it, on Linux sha256sum alarm-recipient-scan.py, on macOS shasum -a 256 alarm-recipient-scan.py, on Windows certutil -hashfile alarm-recipient-scan.py SHA256. A different digest means a different file — not necessarily a hostile one, but not this one.

The repository holds the same file, byte for byte, together with everything needed to re-run the checks this page's claims rest on — so they can be run rather than read about. Issues and pull requests there are read.

Also free

The others

Same idea, a different protocol or a different file. Every free tool.

Next step

Send the recipient that went quiet.

Its status, last failure cause and queued count usually say whether the far end refused the alarm, never saw it, or was never told about it. That read costs nothing either way.