Skip to content

Free tool

poll-scheduler-scan — why a Niagara bus misses the rate it was given

A poll rate on a Niagara network is a budget for the whole bucket, not for each point in it. poll-scheduler-scan reads that arithmetic out of the compiled scheduler with javap and prints it: the three rate defaults, the one-point-per-pass loop, the line that re-schedules a bucket with nothing flooring it, and the point count at which the poll thread stops sleeping altogether. One Python file, read-only, and no station is contacted.

  • Niagara 4
  • driver-rt.jar
  • javap
  • 7 steps
  • Read-only
  • Python 3

Install and run

There is no install. One file, standard library only, and the only thing it needs besides Python is a javap — which it looks for in $JAVAP, on PATH, under $JAVA_HOME, in the JDK Niagara ships, and only then at a Debian path. It reads two compiled classes and contacts nothing.

# download it
curl -O https://plantroomlabs.com/tools/poll-scheduler-scan.py

# read the scheduler in the default install
python3 poll-scheduler-scan.py

# or name the installation to read
python3 poll-scheduler-scan.py /opt/Niagara/Niagara-4.15.5.22

Nothing is written, no station is contacted and no device is polled. It disassembles javax.baja.driver.util.BPollScheduler and BAbstractPollService, then walks the class in seven numbered steps.

The three rates, and what they are a budget for

The three rates, read out of the static initialiser, not the docs:
  fastRate      1000 ms   (min facet 1 ms, showMilliseconds)
  normalRate    5000 ms   (min facet 1 ms, showMilliseconds)
  slowRate     30000 ms   (min facet 1 ms, showMilliseconds)
  pollEnabled lives on BAbstractPollService, with actions
  enable and disable. While it is false the thread sleeps
  1000 ms a turn and polls nothing.

Those are read out of the static initialiser rather than the documentation, which matters because the number a sentence repeats and the number a class holds are two different facts. pollEnabled lives on the service, not the scheduler: while it is false the thread still wakes, sleeps its second, and polls nothing.

One point per pass, and nothing floors the gap

Step 2. pollQueue polls EXACTLY ONE pollable per pass - one
        ArrayList.get and one poll() call - walking the bucket
        round-robin on its own index cursor. A bucket is due when
        nextTicks <= now + 5 ms.

        It then re-schedules itself, and this is the whole of it:
          next = (rate - size * elapsed) / size   i.e. rate/size
                                                  minus elapsed
          nextTicks = now + next
        where size is the bucket's point count and elapsed is how
        long that single poll actually took. An empty bucket
        re-schedules at 1000 ms.

        There is no Math.max in the method. Nothing floors next.

Step 3. So when a poll takes longer than its share of the rate,
        next falls to zero and then below it, nextTicks lands at
        or before now, and the bucket is due again at once. The
        sleep itself is behind an ifle, so zero is already enough.
        computeSleep() starts at 1000 ms and takes the min over all
        three buckets' nextTicks
        minus now, so it returns that same number or less, and
        the sleep is skipped outright.

        The poll thread stops sleeping. It does not log, it does
        not set a status, and it does not slow the bus down. It
        just stops idling, and the real cycle time stretches to
        size * elapsed instead of the rate that was asked for.

        Derived from the constants above, not asserted - the
        largest fast bucket that still leaves a positive gap
        between polls, so the thread still sleeps at all:
            2 ms per poll ->   333 points, and 334 points gives next = 0 ms
            5 ms per poll ->   166 points, and 167 points gives next = 0 ms
           10 ms per poll ->    90 points, and 91 points gives next = 0 ms
           20 ms per poll ->    47 points, and 48 points gives next = 0 ms
           50 ms per poll ->    19 points, and 20 points gives next = 0 ms
        On an RS-485 multidrop a request and its reply is tens of
        milliseconds, so those are the point counts that matter.

The rate is divided, not applied

A bucket re-schedules itself to rate/size minus however long that one poll took. Doubling the points in a bucket halves each point's share of the rate. Nothing about the configuration changed, and nothing warns.

There is no Math.max in the method

When a poll takes longer than its share, the gap goes to zero and then below it, so the bucket is due again at or before now. The sleep sits behind an ifle, so zero is already enough to skip it.

The failure mode is silence, not slowness

The thread does not log, set a status or throttle. It stops idling, and the real cycle time stretches to size * elapsed instead of the rate that was asked for. The point counts in the table above are computed from the constants read in this run, not asserted from experience.

The evidence a station gives you for this

Step 6. What evidence you get, which is 14 properties, all String:
          averagePoll       busyTime          totalPolls
          dibsPolls         fastPolls         normalPolls
          slowPolls         dibsCount         fastCount
          normalCount       slowCount         fastCycleTime
          normalCycleTime   slowCycleTime
        Every one is flags 3 - READONLY|TRANSIENT - and defaults
        to the string "-". statisticsStart defaults to
        BAbsTime.NULL, and resetStatistics is flags 128
        (CONFIRM_REQUIRED).

        Being Strings, they are already rounded and already
        formatted when you read them:
          cycle times  "average = N ms", integer division, and
                       "-" until the first cycle completes
          counts       plain below 10000, then (n/1000) + "k"
          durations    "Nms" below 10000 ms, then (n/1000) + "sec"
        and they are recomputed only every 10000 ms.

        TRANSIENT means they are not saved, so a station restart
        clears them. String means you cannot put a history
        extension on them, cannot link them to a numeric and
        cannot set an alarm on them. The one number that says your
        bus is saturated is a piece of text, read by eye.

Every statistic is a String, flagged read-only and transient, and recomputed once every ten seconds. That is the part worth knowing before you promise a customer a trend: a String cannot carry a history extension, cannot be linked to a numeric and cannot raise an alarm, and transient means a restart clears it. The one number that says the bus is saturated is a piece of text, read by eye.

What follows from it

Three consequences worth designing for
----------------------------------------------------------------------
  1. Overrun is silent and self-inflicted. The scheduler reacts
     to a slow bus by not sleeping, never by reporting. If your
     driver needs an operator to know the cycle has stretched,
     the driver has to say so itself.
  2. Rate is a budget per bucket, not per point. Doubling the
     points in a bucket halves each point's share, so adding
     points to a working network can push it over with no
     configuration change anywhere.
  3. The statistics cannot be trended. If cycle time matters to
     your customers, expose it as a numeric of your own.

Three checks one station settles in an afternoon
----------------------------------------------------------------------
  1. Put N points on one slow bus at the default 1000 ms fast rate
     and read fastCycleTime: it will say "average = ..." well
     above 1000 once N * elapsed exceeds it.
  2. Watch the poll thread's CPU while that is true - the sleep
     is being skipped every pass.
  3. Open a graphic with many points on it and watch the periodic
     poll stall while the dibs stack drains.

What it reads, and the limit of what the output is worth

It reads driver-rt.jar with javap and nothing else. No station is contacted, no point is polled and nothing is written — which is also the limit of what the output is worth. It is a reading of compiled code, so it tells you what the scheduler is built to do rather than what a particular bus did yesterday.

The numbers above came from one installation, 4.15.5.22, and the output names it. The controller this work usually targets runs 4.14.0.162. A different Niagara version is a different answer, so run it against yours rather than trusting this page's figures.

The engineering reasoning this belongs beside is why a station polls too slowly: what to measure on a live station before changing a rate, and the order to change things in.

The file you are downloading

Published here so the download is checkable rather than trusted. Both figures are read off the file served at /tools/poll-scheduler-scan.py when this page is built, so they cannot disagree with it.

PropertyValue
Filepoll-scheduler-scan.py
Size23,627 bytes
SHA-256 86d22bfbb759067505385967af6dc466ce1711b8c4f3a7317ad6a138d57baf2b
Licence MIT — LICENSE.txt
Source github.com/UsamaIqbal0304/poll-scheduler-scan

To check it, on Linux sha256sum poll-scheduler-scan.py, on macOS shasum -a 256 poll-scheduler-scan.py, on Windows certutil -hashfile poll-scheduler-scan.py SHA256. A different digest means a different file — not necessarily a hostile one, but not this one.

The repository holds the same file, byte for byte, together with everything needed to re-run the checks this page's claims rest on — so they can be run rather than read about. Issues and pull requests there are read.

Also free

The others

Same idea, a different protocol or a different file. Every free tool.

Next step

Send the network's point counts and poll rates.

The arithmetic above decides whether a bucket can hold them, and it is arithmetic rather than opinion, so the answer comes back in writing with the numbers it was computed from - whether or not anything gets built afterwards.