Skip to content

Free tool

alarm-route-scan — what a station does to an alarm on the way out

Alarm routing reads like a fan-out: a point goes off-normal, a class routes it, recipients are notified. In the bytecode it is one queue drained by one thread, with a coalesce rule that discards duplicates and an invocation that reports success() for an alarm nobody delivered. alarm-route-scan reads alarm-rt.jar and baja.jar with javap and prints the six findings that follow from that, in the order the station reaches them.

  • Niagara alarms
  • Coalescing
  • Reads the shipped jars
  • No station needed
  • Read-only
  • MIT

Install and run

One file, standard library only, and nothing on the network: it reads two jars on disk. It needs a Niagara installation and a javap on the path, which the JDK that ships with Niagara already provides. Point it at the installation or set NIAGARA_HOME.

# download it next to wherever you are working
curl -O https://plantroomlabs.com/tools/alarm-route-scan.py

# read the alarm service out of an installation
python3 alarm-route-scan.py /opt/Niagara/Niagara-4.15.5.22

What it printed here

This is the whole of one run against the installation on the machine that built this page, pasted by the script that publishes it rather than retyped. Three of the six findings change what an integration can assume.

The invocation that loses a coalesce collision reports that it succeeded. When two routes of the same record sit in the queue together one is coalesced away, and the loser is marked finished with no throwable — so its IFuture.success() returns true. Code that branches on success() cannot tell a delivered alarm from a discarded one, and neither can the caller, because both post methods drop the boolean that enqueue returns.

The property that looks like an on/off switch selects between two rules. coalesceAlarms defaults to true, and in that default the coalesce key is the record UUID, the alarm class and the action — not the source state. Setting it to false switches in the comparison that does include source state. It reads like it turns coalescing on; it chooses which coalescing applies.

Persistence comes before people. Inside doRouteAlarm the alarm database write happens before fireAlarm. A failed write is logged SEVERE and rethrown, so no recipient is notified at all — a full disk on a controller is an alarm outage, not a logging problem.

$ python3 alarm-route-scan.py $NIAGARA_HOME
alarm-route-scan - what a station does to an alarm on its way out
Niagara home: /opt/Niagara/Niagara-4.15.5.22
read from: modules/alarm-rt.jar, modules/baja.jar (javap, no station running)

1. one queue, one thread, and no ceiling

   BAlarmService holds the alarm queue in a field and hands it out through
   fw(601); BAlarmClass.post(routeAlarm, record) asks for it that way and
   enqueues a AlarmClassRouteAlarmInvocation. So every alarm route in the
   station, from every driver and every control point, goes through one
   CoalesceQueue drained by one Worker thread named Alarm:ServiceWorker.
   Workers constructed in serviceStarted(): 1.

   The queue is built with the no-argument constructor, which passes
   maxSize = 2147483647. Queue.enqueue throws QueueFullException only
   when size >= maxSize, so on this build that exception is unreachable
   in practice: a burst the worker cannot keep up with is not refused,
   it is accumulated. Hash table: max(16, min(maxSize/3, 101)) = 101 buckets
   at load factor 0.75.

   The depth is published in exactly one place: the service's spy page,
   as workInAlarmQueue = Queue.size(). It is not a property, so it cannot
   be linked, trended or alarmed on.

2. the coalesce key, and the property name points the other way

   coalesceAlarms: default true, flags 0 (none)

   make() reads it and picks the class:
     coalesceAlarms true  -> AlarmClassRouteAlarmInvocation
     coalesceAlarms false -> CoalesceUuidOnlyInvocation

   Both equals methods compare the alarm class instance, the action and
   the record UUID. Only CoalesceUuidOnlyInvocation also compares the
   record's source state, in equals and in hashCode alike.

   So with the default - coalesceAlarms true - two routes of one record
   that are in the queue at the same time collapse into one whatever
   their states are. Turning the property off is what keeps them apart.

3. the invocation that loses the collision says it succeeded

   CoalesceQueue.enqueue finds the equal entry, calls
   existing.coalesce(incoming), stores the result back into that entry's
   slot - so the survivor inherits the earlier arrival's queue position -
   and returns false.

   coalesce() sets finished = true on the invocation being dropped and
   returns the incoming one. throwable is left null. success() is
   finished and no throwable, so on the dropped invocation it returns
   true: the IFuture says the alarm was routed, and doRouteAlarm never
   ran for it.

   Neither caller can see it either way. BAlarmClass.post and
   BAlarmService.post both pop the boolean enqueue returns.
   (BAlarmService.post only queues escalateAlarms itself; everything
   else goes to the superclass.)

4. the database write comes before the people

   doRouteAlarm, by offset: AlarmDbConnection.append at 169 or .update at
   178, then fireAlarm at 367. An Exception out of that write is
   logged SEVERE 'Cannot write alarm.' and rethrown as an AlarmException, so
   fireAlarm is never reached and no recipient is notified.
   Persistence first, people second.

   A ServiceNotFoundException handler covers the method to offset 503 of
   519 - effectively all of it - and its handler logs and returns. Upstream,
   AlarmSupport logs SEVERE 'Unable to route new alert for' and returns null
   when there is no alarm service at all.

5. escalation is off, and it is cumulative

   escalationLevel<n>Enabled / escalationLevel<n>Delay, by level:

   level 1  enabled false  flags 0  delay  300000 ms  min facet  60000 ms
   level 2  enabled false  flags 0  delay  900000 ms  min facet 120000 ms
   level 3  enabled false  flags 0  delay 1800000 ms  min facet 180000 ms

   escalationTimeTrigger: interval 1 minute(s), flags 4 (HIDDEN)
   escalateAlarms action: flags 2068 (HIDDEN|ASYNC|NO_AUDIT)
   - ASYNC, so the once-a-minute escalation scan is queued onto the same
     single worker thread that delivers alarms.

   Which topic fires is decided by a string facet on the record. doRouteAlarm
   adds 'escalated' with the empty string when it is absent, so a first
   route fires no escalated topic. The tests are cumulative:
     fireEscalatedAlarm1 on level1, level2, level3
     fireEscalatedAlarm2 on level2, level3
     fireEscalatedAlarm3 on level3

   Topic flags as declared:
     alarm            8 (SUMMARY)
     escalatedAlarm1  0 (none)
     escalatedAlarm2  0 (none)
     escalatedAlarm3  0 (none)
   changed() ors SUMMARY into the matching escalated topic when a level is
   enabled and masks it out again when it is disabled.

   Worth a look: the level 2 enable branch reads the flags of
   escalatedAlarm1 and writes them to escalatedAlarm2.
   Levels whose branches agree: 1, 3.

6. ackRequired, and where it is really used

   ackRequired: default 7 = TO_OFFNORMAL|TO_FAULT|TO_NORMAL, flags 0 (none)
   not set by default: TO_ALERT

   doRouteAlarm calls getAckRequired() 1 time(s).
   Its includes(sourceState) result is stored in a local that the
   method never loads again - a dead store. The live consumer is
   AlarmSupport.isAckRequired, called when the record is created.

   AlarmSupport.isAckRequired returns false outright when the alarm
   class reference is null, so a record whose alarm class name does not
   resolve is created with ackRequired false rather than defaulting to
   true.

   BAlarmClass.routeAlarm action flags 20 (HIDDEN|ASYNC)
   the four alarm counts share flags 67 (READONLY|TRANSIENT|DEFAULT_ON_CLONE)
   logger name: alarm

three checks one station settles in an afternoon

   1. Open Services/AlarmService and read coalesceAlarms. If it is true,
      ask whether any driver on the station re-routes one record.
   2. Open the AlarmService spy page and watch workInAlarmQueue while the
      busiest panel or device is made to report a burst. A number that
      does not return to zero is the worker falling behind.
   3. Grep your own alarm code for the IFuture that post(routeAlarm)
      returns. If anything branches on success(), it cannot tell a
      delivered alarm from a coalesced-away one.

Nothing above was measured against a running station: it is read out of
the shipped jars. A reference counted here is not a fault - it is where
to look.

What it reads, and what that does not cover

It reads alarm-rt.jar and baja.jar with javap and nothing else. No station is contacted, no alarm is raised and nothing is written — which is also the limit of what the output is worth. It is a reading of compiled code, so it says what the framework is built to do rather than what a site did last night. The three checks at the end of the output are the ones a single station settles in an afternoon.

The numbers above came from one installation, 4.15.5.22, and the output names it. The controller this work usually targets runs 4.14.0.162. A different Niagara version is a different answer, so run it against yours rather than trusting this page.

Why an alarm never reached anybody is the station side of the same question — the configuration that stops an alarm before any of this runs — and alarm-recipient-scan reads the other end of the path, where the far end is down and the recipient decides what to do about it.

The repository

The same file, MIT licensed, at github.com/UsamaIqbal0304/alarm-route-scan. Its README carries the findings and a run of the output, so the repository is checkable without downloading anything. Issues and pull requests are read.

The file you are downloading

Published here so the download is checkable rather than trusted. Both figures are read off the file served at /tools/alarm-route-scan.py when this page is built, so they cannot disagree with it.

PropertyValue
Filealarm-route-scan.py
Size34,542 bytes
SHA-256 c8a9d01fe65b93f2e3d5d97d3a957337ba71edcdbb736e6851464e16c75ead3c
Licence MIT — LICENSE.txt
Source github.com/UsamaIqbal0304/alarm-route-scan

To check it, on Linux sha256sum alarm-route-scan.py, on macOS shasum -a 256 alarm-route-scan.py, on Windows certutil -hashfile alarm-route-scan.py SHA256. A different digest means a different file — not necessarily a hostile one, but not this one.

The repository holds the same file, byte for byte, together with everything needed to re-run the checks this page's claims rest on — so they can be run rather than read about. Issues and pull requests there are read.

Also free

The others

Same idea, a different protocol or a different file. Every free tool.

Next step

Send the alarm that nobody got.

The queue, the coalesce rule and the recipient are three different places an alarm can stop, and the station already holds enough to say which one it was. That read costs nothing either way.