Skip to content

Note

Why a Modbus network sits at 95% busy time

A Modbus network reporting 95% busy time is not short of CPU, and adding cores will not move it. It is one thread spending almost all of its life waiting on sockets, and the fix is arithmetic rather than hardware.

  • Modbus
  • Performance
  • Station engineering

Written September 2026

What the number is measuring

The poll scheduler brackets each poll with a clock read, accumulates the result into a running total, and every ten seconds prints that total over its own uptime as a percentage. So busy time is the duty cycle of the poll loop — the share of wall-clock time the scheduler spent inside a poll — and nothing else. It is not host CPU, it is not the JVM, and it is not shared with any other network.

That changes what a high number means. At 95%, the poll thread is nearly always inside a transaction. No amount of extra vCPU helps, because nothing else on the machine runs that loop, and the thread is not computing anything while it is busy.

One thread per network, and that is the whole lever

Starting a poll scheduler creates exactly one plain Java thread, named Poll: followed by the network's name. The scheduler is a property of the network component, so every additional Modbus TCP network in the tree is another scheduler and another thread.

Nothing else about the configuration changes that. The TCP port is irrelevant: each Modbus TCP device carries its own address, port and socket, and opens its own connection regardless of which network it sits under. Splitting devices across several networks is the mechanism for parallelism — it is not a hint or a tuning parameter.

Why splitting works at all. The poll thread does the wire I/O itself. A poll group's read calls straight through to the device's register read, which blocks on the socket until the answer arrives or the timeout expires. So the thread's busy time is mostly socket wait, and two threads wait in parallel where one waited in series.

Timeouts dominate the number

The driver framework defaults responseTimeout to 500 ms and retryCount to 1. A device that has been switched off, re-addressed or firewalled therefore costs a full second of the cycle every time its turn comes round, and the poll group catches the exception and logs a poll device timeout at info level — which is not on by default and does not raise anything.

Put that next to a healthy transaction on a switched LAN, which is a few milliseconds to a few tens of milliseconds. One silent device can be worth fifty working ones in the busy-time figure, and it does so without appearing anywhere an operator looks. Before rebalancing anything, find the devices that are not answering.

The rate belongs to the poll config entry, not the point

A Modbus proxy point does not get polled on its own. Its address is looked up against the device's active poll config entries for that register type, and every point that falls inside the same entry shares one poll group. A point matching no entry becomes its own group and is polled individually, which is the pointPoll data source visible on the proxy extension.

The group's rate is then decided by a single rule: start at Slow, walk the subscribed points in the group, and keep the fastest one found. Fast beats Normal beats Slow. So one Fast point inside an entry covering fifty consecutive registers puts all fifty registers in the Fast bucket.

Two consequences follow. To give one register its own rate you must split the poll config entry, not change the point. And a point with no link, no history extension and nobody viewing it is not subscribed, so it does not drag its group's rate while it sits idle.

Splitting an entry trades registers for round trips. One entry is one request for up to 125 registers. Split that fifty-register block into three so that one register can run Fast, and the device now answers one small fast request plus two slow ones. Worth it at 1 second against 30. Close to pointless when the two rates sit near each other.

The arithmetic that tells you how many networks you need

Each rate is a cycle target, not an interval. The scheduler polls one group per pass from a bucket, then sleeps for the bucket's rate minus the work just done, divided by the number of groups in it — it is trying to walk the whole bucket once per rate. Add groups and every existing group's share of the second shrinks. When the sum of poll times exceeds the rate, that sleep goes negative and the bucket free-runs at full speed, which is the 95% reading.

So the condition for a bucket that keeps up is simply:

QuantityWhere it comes fromExample
Groups in the bucket Poll config entries across the network's devices, plus any unmatched points 250
Time per transaction Average Poll on the scheduler, once dead devices are removed 20 ms
Cycle needed Groups multiplied by time per transaction 5 s
Rate configured Fast, Normal or Slow on the scheduler 5 s

That example is exactly saturated: one slow device, one retry, and it is over. Halve the groups per network and the same work sits at 50% with room for a bad afternoon. This is why splitting helps in practice and also why it stops helping — the win is per-network group count, so moving two devices does nothing and moving half the estate does.

Read the scheduler's own figures rather than guessing at them. Average Poll gives the time per transaction directly, the per-bucket counts give the group numbers, and the cycle times give the answer the arithmetic was predicting. Reset the statistics before and after a change so the comparison is between two measurements.

What to change, in order

  • Find the devices that are not answering. Each one is worth roughly fifty healthy transactions. Turn on the network's log long enough to catch the timeouts, or read each device's status.
  • Move points out of Fast. Remember the rule is per entry and fastest-wins, so one forgotten point is enough to hold a whole block up there.
  • Check the poll config entries themselves. Contiguous blocks read in one request are the cheapest thing Modbus can do; a device whose points all fell through to pointPoll is paying one round trip per point.
  • Then split the networks, in halves rather than by twos, and measure after each split.
  • Offset the rates. Using intervals that do not divide into each other keeps the buckets from lining up on the same second.

What is not on that list is the thread pool. The poll thread is created directly by the scheduler, not drawn from a pool, so pool sizing does not touch this number.

The companion note on poll rates and tuning policies covers how a point ends up in a bucket across drivers generally, and Modbus register addressing covers the entry addresses themselves. Where a bus has to be rebuilt rather than retuned, protocol integration is that work, and station engineering is the station around it.

Related

Where this comes up in the work

More notes

Other things worth writing down

Next step

Tell us the version, the hardware, and what it has to do.

You will get a written scope and a fixed price against it. If the honest answer is that you do not need us, you will get that instead.