· Engineering  · 61 min read

Five Doors Into One Siemens Controller: What Each One Costs You

How you ask a Siemens PLC for data matters more than which protocol you use. Reading the same hundred values off the same controller, a client that groups its requests is a hundred times faster on Modbus and five to fifteen times faster on S7Comm. That is larger than any difference between the protocols themselves. A measured look at what protocol and driver choice costs a PLC, a network and your budget.

How you ask a Siemens PLC for data matters more than which protocol you use. Reading the same hundred values off the same controller, a client that groups its requests is a hundred times faster on Modbus and five to fifteen times faster on S7Comm. That is larger than any difference between the protocols themselves. A measured look at what protocol and driver choice costs a PLC, a network and your budget.

1. The pitch, and the bill

Walk any trade fair and you hear the same answer. Standardise on OPC UA and talking to machines becomes a solved problem.

Most of that is true. OPC UA really is a standard, really is vendor-neutral, and really does carry an information model the older protocols do not have. For many plants it is the right choice.

What the pitch leaves out is the bill, and it arrives in one of two ways.

The first is a rollout. A brownfield line that has run for fifteen years is told to standardise. Nobody re-specifies the controllers, because the controllers work fine. Then the network carries several times the traffic it used to, and the CPUs spend part of every cycle on communication they were never sized for.

The second is slower, and much harder to see coming. An Industry 4.0 project starts small and works well. So it grows. Somebody adds a dashboard. A quality system wants the same values, faster. Maintenance wants trends. Someone wants all of it archived. Every request is reasonable on its own, and every one gets granted, because the last one was fine.

Nobody ever decides to overload a controller, or a network. You arrive there one good idea at a time, until the PLC can no longer serve the information and run the machine at the same time.

Both roads end in the same place. The line stops running properly. Not dramatically. It just gets less reliable, in ways that take a long time to trace back to a connectivity decision.

So this article asks a narrow question. Not "which protocol is best", but what does it cost to read data out of a controller you already own — and how much of that cost is the protocol, how much is the controller, and how much is your own client.

Other vendors get the same treatment in their own articles, starting with Beckhoff’s TwinCAT controllers. There will also be one that crosses vendor lines on purpose.

2. How I measured it

Five Siemens controllers. Six ways in. One number that changes: how many values I ask for at once.

2.1. One run

A run is one controller, one protocol, one size. It opens a connection, then reads all the values together, twenty times over, with the order shuffled each round.

Every figure in this article is one of those reads. Each row reports the average of its twenty, with the two fastest and two slowest dropped, so one slow round cannot move it.

The connection is timed separately and is not in the result. Nothing is ever written: this is a read benchmark, because reading is what a shopfloor does almost all of the time.

The sizes are 10, 25, 50, 100, 200, 500, 1000, 2000 and 3000 values. Where a controller or a protocol cannot manage a size, that row is left out rather than quietly shrunk, so a row always says what it really read.

2.2. The values

Every value is one element of the same array, and I always read elements 0 to N-1. They sit next to each other in memory.

That is deliberate, and it is the best case for one of the things being measured. More on that below, because it matters.

They are also all the same simple type, and that is deliberate too.

An earlier round of this work varied the opposite thing. Not how many values, but what kind. Nine types that every controller and every protocol here can handle. Then twenty-five that most can, which already leaves out Modbus and the oldest controller. Then fifty-six involving structure members, nested structures and single cells of a matrix, which only the protocols that address by name can express at all.

That answered a different and equally real question — what each protocol can reach — and those findings still stand. S7Comm cannot name a structure member, whether you ask it for ten of them or three thousand. But it makes a poor curve, because each set runs on a different subset of the bench, so the rows are not comparable as the set changes.

So this curve holds the type constant at something every device supports, and varies only the count. Capability is a separate axis, and mixing the two is how you end up comparing three different machines and calling it a trend.

2.3. Three things that make the numbers trustworthy

A control that should not move. Every measurement has a twin that reads values spaced far enough apart that no client can merge them. If the gains came from anything other than merging, the twins would move too. They agree within 5% on every controller.

The connection is timed separately. It used to be inside the result, which quietly flattened every small measurement, because both halves of a comparison paid the same fixed cost. On a nine-value run from an earlier sweep, Modbus read 0.9x that way and 6.2x with the connection taken out — the same measurement, and one of those numbers says merging does not work.

A discarded warm-up run. The first merged run in a process pays to compile the merging code: 1428 ms cold, 330 ms warm, while its twin moved 2%. Without that warm-up the benchmark reported a 2.5x slowdown that does not exist. I published that slowdown once. It was my own JVM.

2.4. How to read the tables

Rows that share a controller compare protocols. Rows that share a protocol compare controllers. Change both at once and the row compares nothing you can name.

3. A read costs one request, not one byte

Start here, because everything else follows from it.

Over classic S7Comm I read one contiguous block of memory and varied only how many bytes it held:

Table 1. One S7Comm read of a contiguous block. Mean time.
Bytes readS7-1200 G2Old S7-1212C

1

2.07 ms

12.62 ms

10

1.99 ms

12.60 ms

100

1.97 ms

12.62 ms

200

2.06 ms

12.95 ms

Flat. Two hundred times the data for the same time.

The tables further down say the same thing from the other direction. On one controller, reading a hundred values takes 1.9 ms in one request and 318 ms in a hundred requests — for ten times fewer bytes in the slow case.

A request to a PLC is almost all fixed cost. The controller has to notice the request, schedule it around the program it is running, build an answer and send it. The data itself is nearly free.

So the way to make a client faster is not to ask for less. It is to ask fewer times.

4. The curve

Which raises the obvious question: how much fewer, and when does it stop helping?

I ran the same values at nine sizes, with the client merging requests and without, on four controllers. Every figure below is one batched read against one batched read — the same hundred or thousand values, asked for once each way.

4.1. On Modbus it is enormous

Modbus has no way to ask for several things at once. One request is one contiguous range of registers and nothing else. So a client that does not merge sends one request per value, and a client that does packs up to 125 registers into each.

Table 2. How many times faster one merged read is than the same read unmerged. Modbus TCP.
ValuesOld S7-1212CS7-1212C G2S7-1511S7-1516Requests

10

14.1x

10.5x

10.3x

10.3x

10 → 1

25

25.6x

25.5x

25.6x

25.8x

25 → 1

50

59.7x

52.8x

50.5x

51.8x

50 → 1

100

97.4x

127.4x

107.8x

103.3x

100 → 1

200

99.5x

109.8x

106.7x

103.3x

200 → 2

500

—

126.5x

122.6x

130.4x

500 → 4

1000

—

125.5x

118.9x

—

1000 → 8

S7-1212C G2S7-1511S7-1516S7-1212C (old)0x25x50x75x100x125x10255010020050010003000values per read (log scale)G2151115161212C
Modbus TCP: how many times faster one merged read is than the same read unmerged.

The last column is the whole explanation. Up to a hundred values the merged version is one request, so the ratio is simply how many requests you avoided. After that the 125-register limit starts to bite and the merged version needs several too, so the ratio stops climbing.

The gaps are rows where the unmerged read never finished. More on that below.

4.2. On S7Comm it is smaller, and it depends on the controller

S7Comm can carry many items in one request. So an unmerged client still gets its values without one request per value — but each item costs a twelve-byte header inside the request, so the request fills up and has to be split.

How many splits depends on the negotiated request size, and that is what separates the controllers.

Table 3. How many times faster one merged read is than the same read unmerged. S7Comm.
ValuesOld S7-1212C (240 B)S7-1212C G2 (960 B)S7-1511 (960 B)S7-1516 (960 B)Requests, 240 BRequests, 960 B

10

1.9x

1.5x

1.7x

1.8x

1 → 1

1 → 1

25

4.3x

1.9x

3.0x

2.5x

1 → 2

1 → 1

50

7.5x

2.9x

4.4x

3.8x

1 → 3

1 → 1

100

15.2x

5.3x

10.0x

6.6x

1 → 6

1 → 2

200

14.9x

13.1x

18.7x

12.3x

2 → 11

1 → 3

500

14.5x

18.1x

19.7x

13.8x

5 → 27

2 → 7

1000

14.0x

20.8x

25.3x

16.2x

10 → 53

3 → 13

2000

13.7x

23.4x

28.6x

18.1x

20 → 106

5 → 26

3000

14.2x

23.4x

29.9x

20.0x

29 → 158

7 → 38

Two request columns, because S7Comm needs two. Modbus sends one request per value on every controller, so one column describes them all. S7Comm’s count depends on how much it can fit in a request, and the negotiated size is not the same on every device: 960 bytes on three of these, 240 on the old S7-1212C. Same protocol, same values, different arithmetic.

Which is what the ratio columns are really showing. At a hundred values the small-PDU controller goes from six requests to one and gains 15.2x; the large-PDU ones go from two to one and gain 5.3x to 10x. The worse your request size, the more merging is worth to you.

S7-1212C G2S7-1511S7-1516S7-1212C (old)0x5x10x15x20x25x30x10255010020050010003000values per read (log scale)G2151115161212C
Classic S7Comm: the same comparison. Note the y-axis — this is a quarter of the scale above.

At three thousand values an unmerged read takes 38 requests on the 960-byte controllers and 158 on the 240-byte one. Merged, it takes 7 and 29.

4.3. What plateaus is the improvement, not the protocol

Read those two tables carefully, because they are easy to misread — I misread them myself first.

On Modbus the ratio stops climbing at about a hundred values. On S7Comm it keeps climbing much further: the small-request controller settles around fifteen from a hundred values on, while the larger ones are still improving at three thousand.

But a plateau in the ratio is not a reason to stop batching. If you read ten thousand values, Modbus still gives you its hundredfold and S7Comm still gives you its twenty or thirty. The saving does not go away. It simply stops getting better.

So there is no need to chop a large read into chunks of a hundred. It would not hurt much, and it would not help either. Ask for what you need in as few requests as your client will build, and the curve says you will be at or near the best available.

The two tables also say something you can act on without measuring anything:

How much grouping is worth to you depends on whether your protocol can already carry several items in one request. If it cannot, a client that does not group is throwing away two orders of magnitude.

5. Where your values sit

Merging works by reading a span of memory in one go. If your values are neighbors, that span is nearly all data you asked for.

If they are scattered across a data block with big gaps between them, the client is reading bytes nobody wants in order to save round trips. At some point that trade stops paying, and if they are far enough apart it buys nothing at all.

Every number above reads one contiguous array, which is the best case for merging. On your plant it will depend on your layout.

For OPC UA the same array is the worst case. Its server treats every element as a separate node read through an index range, and on the G2 it answered separate variables about eight times as fast as array elements. So every OPC UA row in this article up to the next section is a lower bound. The section after the six-way table measures OPC UA the way most projects are written.

Modbus on a SIMATIC is unusually well placed here. You cannot speak Modbus to an S7 at all without building a data block for it and mapping your values into register ranges by hand, so they end up beside each other because the protocol made you put them there. The layout that makes merging work is the layout you were forced into anyway.

6. Six ways in, side by side

Now the protocols against each other. One controller, one read carrying a hundred values:

Table 4. One read of a hundred values on an S7-1212C G2. Merged rows use the client’s optimizer; plain rows are the protocol on its own. Connect is opening the connection: a TCP socket, plus whatever the protocol negotiates on top of it.
How you askOne readRequestsBytesConnect

S7Comm, merged

1.9 ms

1

256 B

69 ms

Modbus, merged

2.5 ms

1

221 B

75 ms

S7Comm

10.0 ms

2

1880 B

71 ms

S7CommPlus

22.0 ms

8

2736 B

572 ms

S7 Web API

188 ms

—

14200 B

421 ms

OPC UA

212 ms

1

4730 B

2048 ms

OPC UA, secured

214 ms

1

4800 B

9175 ms

Modbus

318 ms

100

2300 B

82 ms

Look at the two Modbus rows before anything else. Same controller, same hundred values, same protocol. One of them asks once and takes 2.5 ms. The other asks a hundred times and takes 318. That is the whole argument of this article in two rows.

One gap to explain before the rest: the web interface has no entry in the requests column. Its packets carry TLS handshakes and renewed sockets as well as the read itself, so there is no honest way to attribute a count to one read. The dash means not separable, not none.

Six things in that table.

S7Comm is the cheapest thing a SIMATIC can do, by a wide margin, and it ships with the controller. It is also unauthenticated, unencrypted and blind to variable names: it reads bytes at an address, so your client has to know the memory layout from somewhere else. Move a variable in your program and something breaks quietly, with no compiler to catch it.

S7CommPlus costs about eleven times the merged S7Comm read and buys two real things. Variable names instead of addresses, and a secured session. It is free and already on the controller.

Read as array elements, OPC UA costs a hundred times the merged S7Comm read and ten times S7CommPlus. Array elements are its worst case. The next section measures the better one. Note the requests column while you are there: OPC UA reads all hundred values in a single request. Its cost is not round trips. It is what the server does with them.

The connect column is where the protocols differ most, and it is a ladder. Modbus is the only one with no handshake at all. There is no session to establish because Modbus has no concept of one, so its 82 ms is a TCP socket and nothing else. S7Comm does negotiate, but barely — a hundred bytes or so each way, and that exchange is where the request size discussed further down gets agreed. S7CommPlus sets up a proper session and costs half a second. OPC UA costs two seconds, and nine with security switched on.

You pay that once per connection. So it matters enormously if your client reconnects for every poll, and not at all if it holds the connection open.

The bytes column does not predict the time column. OPC UA moves 4730 bytes in one request and takes 212 ms. Unmerged Modbus moves half that and takes 318 ms across a hundred requests. The web interface moves three times as much as OPC UA and is faster. Whatever you are paying for, it is not data volume.

And the same weakness that makes S7Comm blind is what makes it tunable. Because an S7Comm address is a position rather than a name, a client can see that your values are neighbors and fetch them in one span. A protocol that uses names cannot: arrInt[7] and arrInt[8] are two opaque names, and nothing in the request says they sit next to each other. That is why only S7Comm and Modbus have a merged row. The cheapest door is the one with the most headroom left in the client.

6.1. OPC UA over separate variables

Most projects publish separate variables over OPC UA, not the elements of one array. So the next run reads that shape, one node per value, and puts it beside S7Comm and S7CommPlus at the same counts. Time for one read:

Table 5. One read, OPC UA over separate variables.
CPUValuesS7Comm, mergedS7CommS7CommPlusOPC UA

S7-1212C G2

80

3.3 ms

11 ms

19 ms

25 ms

S7-1212C G2

500

6.2 ms

52 ms

125 ms

129 ms

S7-1212C G2

1000

18 ms

93 ms

295 ms

254 ms

S7-1511

80

4.8 ms

31 ms

57 ms

432 ms

S7-1511

500

12 ms

173 ms

368 ms

2 436 ms

S7-1511

1000

27 ms

341 ms

765 ms

4 411 ms

Old S7-1212C

500

70 ms

929 ms

1 298 ms

3 594 ms (array elements)

Old S7-1212C

1000

147 ms

1 837 ms

2 623 ms

7 137 ms (array elements)

S7Comm is still the fastest on every CPU. Merged, it is 6 to 30 times faster than S7CommPlus. Even unmerged it is 1.4 to 3 times faster.

On the G2, S7CommPlus and OPC UA are level. At a thousand values OPC UA is slightly ahead. The hundredfold gap in the six-way table was mostly the array.

On the S7-1511, OPC UA is the slowest protocol by far, six to eight times behind S7CommPlus. Its server spends about 4.4 ms per variable, whatever the variables look like. Separate variables help there too, but only by about two and a half times.

There is no OPC UA figure at 2500 values. The G2’s server can publish at most 2000 nodes, and about 800 were already taken. The S7-1511 refuses a single read of 2500 items. The old S7-1212C had no memory left for a block of separate variables at all, so its OPC UA rows read array elements.

7. The controller matters more than its age

The four controllers are not equally fast, and the reason is not what people assume.

One takes six times longer per request than another of the same family. The old S7-1212C needs about 12.6 ms for a request the G2 answers in 2.0 ms. Both are S7-1200s.

And it negotiates a 240-byte request where the others manage 960. A smaller request holds fewer values, so the same read needs more of them.

Those two facts explain every slow row on that controller. Its idle cycle does not: it is twice the G2’s, but a read takes four to twenty times as long. Its age does not come into it, except that both facts are things the newer generation improved.

It is worth being careful in the other direction too. The S7-1511 on this bench is three to six times slower per request than the S7-1516 over S7Comm, OPC UA and the web interface. It is the smallest S7-1500 and the S7-1516 is a much faster class, which is the likely explanation but not a proven one. Read its absolute times as that controller’s, not as "an S7-1500’s".

7.1. The setting that is worth a factor of two

A Siemens PLC reserves a share of its cycle for answering questions. In TIA it is Cycle load due to communication.

Same controller, same program, same values:

S7-1511, communication capS7Comm, merged

20%

1511 ms

50%

737 ms

Twice as fast, from a setting.

And the default is not the same on every family. Only the S7-1500 ships with 50%. The S7-300, the S7-400 and the S7-1200 all ship with 20%.

The S7-1200 is the interesting one there. It is not an older controller. It is the S7-1500’s sibling, the low-cost variant of the same generation. So this is not old hardware against new. Siemens give the more capable controller two and a half times the communication budget and leave the less capable one at 20%. The allowance tracks how much headroom a controller has, not how much it needs.

Two consequences. Any figure you quote has to say what the cap was. And before you compare two controllers in your own plant, check they are set the same, or you are measuring the setting.

7.2. What does not cost anything

I expected the PLC’s own program to matter. It does not, at least not here.

On the controller with the busiest program on the bench — an OPC UA server block, a Modbus server block and a model railway’s worth of inert logic — switching all three off changed nothing measurable. Every protocol landed within 0.6% of where it was.

That is measured rather than assumed, and it is worth knowing because "the PLC is busy" is the first thing everyone blames.

One honest qualification. Switching off the OPC UA block only skipped a function call. The server itself lives in the device configuration and kept answering, so this does not show that an OPC UA server is free.

8. What the traffic costs the controller

Everything so far measures what your client experiences. There is a second question, and on a busy plant it is the one that decides whether you have a problem: what does all this asking cost the PLC?

A PLC is not a server. It runs your program in a loop and answers questions in the gaps. Ask too much and the loop gets longer.

A longer loop does not look slow. Decisions are taken once per scan, so stretching the scan does not make the machine slower in any way an operator would call slow. It makes it less precise. The cut lands a millimetre off. The actuator overshoots, sometimes. The reject rate creeps up.

That is close to the worst diagnostic signature a fault can have. It is intermittent. It looks mechanical. Everything a maintenance engineer would sensibly check sits downstream of the cause. Nobody walks up to a machine cutting slightly out of tolerance and asks what protocol the historian is using.

This is also where the second story from the beginning of this article ends up. Nobody adds the dashboard that breaks the line. They add the eleventh one.

8.1. Measuring it from inside the CPU

Every number so far was taken flat out: ask again the moment the answer arrives. That is the right way to compare protocols and the wrong way to ask this question, because flat out saturates the machine and everything converges on the same ceiling.

So this needs a different test. A small block in the PLC program counts its scans and sums their durations, so every rate yields the exact mean cycle under that load. The client polls 80, 500 or 2500 values at a held rate, from once every 2 s to every 10 ms, three passes each, against an idle baseline taken the same way. A positive control proves the figure can move before a flat row is read as "no cost".

The cost is stated as the communication share: how much longer the mean cycle got, as a share of the loaded cycle. A workload that reaches the communication cap from earlier stops growing the cycle and starts waiting, so a cell at the cap measured the limit, not the load. Every run assumed a 20% cap, the S7-1200 default. The S7-1500s ship with 50%, but none of their cells reached even 20%.

Table 6. Largest communication share any rate reached. Four CPUs, every protocol each one offers.
ProtocolOld S7-1212CG2S7-1511S7-1516

S7Comm, merged

at the cap

1%

1%

1%

S7Comm

at the cap

1%

1%

3%

S7CommPlus

at the cap

2%

1%

—

OPC UA

at the cap

1%

0%

—

Modbus TCP

at the cap

0%

17%

0%

S7 Web API

—

2%

1%

—

Table 7. Communication share on the old S7-1212C, 80 values per poll.
Poll everyS7Comm, mergedS7CommS7CommPlusOPC UAModbus TCP

2 s

idle

idle

idle

3%

idle

500 ms

idle

4%

7%

at the cap

idle

200 ms

idle

18%

at the cap

at the cap

idle

50 ms

6%

at the cap

at the cap

at the cap

9%

20 ms

16%

at the cap

at the cap

at the cap

at the cap

10 ms

at the cap

at the cap

at the cap

at the cap

at the cap

Four things come out of that.

Only one of these controllers has a load curve at all. On the G2, the S7-1511 and the S7-1516, no protocol cost the cycle more than 3% at any rate this bench could drive. There is one exception we are still checking: Modbus on the S7-1511, at 17%, where the other two serve the same workload for almost nothing. If your hardware is current and your client is sensible, this is probably not your problem. Better to know that before you go looking for it.

Asking badly costs more than the protocol does. Merged, that controller polls every 200 ms without the cycle noticing. Unmerged, it already spends 18% there. Same protocol, same hardware, same eighty values. The only difference is how the client asked.

The heavier protocols reach the cap first. On the old S7-1212C, OPC UA costs 3% already at one read every 2 s and hits the cap at one every 500 ms. S7CommPlus hits it at 200 ms. At 2500 values only merged S7Comm stays clearly under the cap, up to one read every 500 ms.

And at the cap it stops. That is the communication cap from earlier, doing exactly what it is for. Ask faster and you do not get more data. You get the same data and a stretched cycle.

That last one is worth dwelling on, because it makes the failure quiet. A controller at its cap does not refuse anything. It just stops keeping up, and the first place that shows is the machine rather than the monitoring.

8.2. Subscriptions

OPC UA subscriptions were measured on the G2 and the S7-1511, at 40, 500 and 1000 values. S7CommPlus offers subscriptions too, but our client could not yet establish them on these CPUs. The old S7-1212C’s server does not accept subscriptions on array elements.

Watching values that do not change costs next to nothing. At most 4 µs on the S7-1511 for 1000 values at any requested rate, and nothing measurable on the G2.

Values that change every cycle cost real time. On the S7-1511, 40 of them cost about 30 µs, 3% of the cycle. 500 values of which 100 change every cycle cost about 200 µs, 15 to 17%, about half of what the old S7-1212C spends at its cap. A thousand such values could not be served in this run: a read of them timed out.

Asked for a fixed rate, both servers reported changes only, as Beckhoff’s OPC UA server did in the TwinCAT article. And the S7-1511 delivered at most one value per variable per second, whatever rate was asked for, down to 10 ms.

9. Security is cheap to use and expensive to start

This surprised me, and it is the most useful thing in the article for anyone currently arguing about it.

Look back at the six-way table. One read of a hundred values over OPC UA takes 212 ms unsecured and 214 ms secured. That is a difference of 0.9%.

Push it further. On one controller I timed a single batched read at several sizes:

Values in one readUnsecuredSecured

600

7.47 s

7.48 s

700

8.71 s

8.71 s

1000

12.40 s

12.47 s

Encryption costs about a millisecond per read at 700 values, and 70 ms at a thousand. It is too small to plan around.

The cost is all in the handshake, and that is where it hurts. Opening the connection is a separate number, and it is not small:

Table 8. Time to open one connection. The same client, the same controllers.
How you connectS7-1212C G2S7-1511Old S7-1212C

Modbus

82 ms

62 ms

84 ms

S7Comm

71 ms

78 ms

149 ms

S7 Web API

421 ms

908 ms

4264 ms

S7CommPlus

572 ms

1153 ms

10086 ms

OPC UA

2048 ms

211 ms

444 ms

OPC UA, secured

9175 ms

5819 ms

28472 ms

Read the bottom row twice. Turning OPC UA’s security on takes the connection from two seconds to nine on one controller, and from under half a second to twenty-eight and a half seconds on the old one. That is not a rounding error on a startup cost. That is a client that looks broken.

And S7CommPlus is in the same territory on that controller — ten seconds to open a session, against 149 ms for S7Comm on the same wire.

Which gives a rule you can act on:

Hold your connections open. If security is making your system unusable, the problem is almost certainly your connection handling and not your policy.

That is not a small ask in practice. Plenty of tooling connects, reads and disconnects on a schedule, because that is the simple thing to write, and on that old controller a secured OPC UA poll built that way spends twenty-eight seconds connecting to do 214 ms of work. Every time.

Which is where people reach for the obvious fix and turn security off. Then they are paying a licence fee for a protocol that is now protecting nothing, and the thing that was actually wrong was never the encryption.

9.1. "It uses TLS" and "it is secure" are different claims

Both S7CommPlus and OPC UA can encrypt the channel, and both can demand a credential. So the lazy version of this comparison, where one is protected and the other is not, is simply wrong.

Where the identity lives. On an S7-1500 you set a protection level in TIA and give it a password. That gates the controller, and it is one gate for everyone who connects. OPC UA does separate two things that S7CommPlus runs together: the security of the channel, and who you are. You can require certificates to open the connection and a named user to use it.

But do not mistake authentication for authorisation. On most controllers in the field, getting in is all you get. There is no way over OPC UA to admit one user for reading and another for reading and writing. Everyone who passes the gate has the same access.

That is starting to change. With S7-1500 / ET 200SP / ET 200pro ≥ FW V3.1 (TIA ≥ V19) can grant permissions per user. Michael Grollmus walks through the setup in his video. If your controllers are that new, it is worth the effort. Anything older is still one gate with the same access for everyone.

On those older controllers, the one door that can tell two users apart is the web interface. You can give one account read access and another read and write.

It does not go further than that. There is no way to say that this user may write these tags and not those. But until recently, read-only against read-write was a distinction none of the other doors could make at all. That made it an odd place for it to live: the one least suited to moving data was the only one that could express it.

Whether anyone else can check it. OPC UA is written to a public specification, so you can state your posture in terms an auditor recognises and a second tool can verify. S7CommPlus is a Siemens mechanism. It has been publicly reverse-engineered, and Siemens have their own bulletin on that, but it is not a standard you can point anyone at.

And on OPC UA all of it is optional. It was off by default on two of my controllers. "We use OPC UA" is not a statement about security.

10. What actually stopped, and why

Several rows in my sweep failed. Almost none of them are what they look like.

Nearly every failure was my own timeout running out. The client gives a read request a deadline when it is submitted. An unmerged Modbus read of 1000 values is a thousand requests against one deadline, so once the clock passes five seconds every remaining value fails at once. Raise the timeout to sixty seconds and those rows pass. That is a speed difference crossing a default, not a protocol refusing to do something.

One controller’s OPC UA server does not get cheaper per value as the batch grows. It needs about 12.4 ms per value whether you ask for six hundred or a thousand. So it never refuses; it simply takes longer than whatever timeout the client set. Where it "stops" is wherever you put your deadline.

There was exactly one real size limit in the whole sweep. The web interface on one controller refuses a request body over 64 KiB, which a thousand reads exceeds. Bisected to the byte. Our driver had assumed twice that for a controller that reports nothing, and it now assumes the smaller number.

The general lesson is worth more than any of the specifics. A failure in a benchmark is evidence of something, and the something is usually your own configuration.

11. The second bill

There is a cost that never appears in a benchmark, and on a plant-sized rollout it often decides the argument.

S7Comm, S7CommPlus and the web interface are all on board and free. Nothing to order, nothing to renew.

OPC UA is a separately licensed product. They are runtime licences: Basic on an S7-1200, and Small, Medium or Large on an S7-1500, picked by how big the CPU is.

Nothing enforces it. There is no key to install and no check on the controller. TIA tells you which size your configuration needs, and the CPU serves OPC UA whether you bought one or not. So it is a legal obligation rather than a technical gate, which makes it the kind of cost that surfaces in an audit rather than at commissioning.

Siemens do not publish a price list, so treat these as orders of magnitude from public distributor listings, converted to euros at an approximate rate:

Table 9. Indicative runtime licence cost per CPU
LicenceCoversRoughly

S7-1200 Basic

S7-1200

EUR 60 – 95

S7-1500 Small

CPUs up to 1513

around EUR 100

S7-1500 Medium

CPUs up to 1516

around EUR 320

S7-1500 Large

all S7-1500 CPUs

around EUR 400

Per controller. Twenty controllers is a four-figure sum. Two hundred is a five-figure one.

And there is a loop worth naming, because it is invisible in a feature comparison. The price is set by how big the CPU is, not by how much OPC UA you use:

Choosing the heavier protocol can force a larger CPU, and a larger CPU costs more to licence for the heavier protocol.

Decide the protocol and the hardware together. Deciding the protocol afterwards is how a project finds a line item nobody budgeted for.

11.1. And the CPU decides how much you can publish

There is a harder version of that trap, and it has nothing to do with speed.

Over OPC UA, how many data points a controller can publish is fixed by the CPU, and TIA checks it when the project compiles. Go over and the project does not build. You do not find this out on the shopfloor. You find it out at your desk, after the hardware is bought.

Table 10. Nodes allowed in a user-defined OPC UA server interface. Established by compiling in TIA.
ControllerAllowed

S7-1212C

2000 nodes

S7-1212C G2

2000 nodes

S7-1511

1000 nodes

An array counts as one node per element. So a single array of three thousand values does not fit any of them.

Yes, the S7-1500 number is the smaller one. It also has a way out that the 1200s appear not to have. Through its standard SIMATIC interface an array is one node, read by index range, so it costs one node however long it is — ten thousand elements compile there without complaint. Whether an S7-1200 offers that route at all is an open question, and our evidence says it does not.

If that holds, an S7-1200 cannot publish more than two thousand OPC UA data points by any route.

So picture a simple machine that needs a small CPU and has one large array to publish. Over OPC UA you may end up buying a bigger controller, not because the small one is too slow, but because your project will not compile on it.

That same S7-1200 serves the same three thousand elements over S7CommPlus, in one request, with nothing configured on the PLC. And at the one size both protocols can reach, S7CommPlus reads the array 2.7 times faster as well. Published as separate variables, the G2’s two are level, but then the node limit is the one you are spending.

12. So what should you do?

Ask for what you need in as few requests as your client will build. This is the single highest-value change available to most installations and it costs nothing.

Check whether your client merges adjacent addresses, and on which protocols. If you are on Modbus and it does not, you are leaving two orders of magnitude on the table. If you are on S7Comm you are leaving somewhere between five and thirty times, depending on the controller.

Hold your connections open. Security is nearly free once connected and expensive to establish. A client that reconnects per poll pays the worst of both.

Before you buy OPC UA for variable names on Siemens hardware, price S7CommPlus. It is on the controller and it costs nothing. On the G2 it is about as fast as OPC UA over separate variables. On the S7-1511 it is six to eight times faster. Buy OPC UA for vendor-neutrality, for the information model, or for a security story an auditor recognises. Do not buy it for naming you already have.

If your read volume is modest and your security requirements are real, buy OPC UA. The overhead is affordable at low rates and you get certificates, user authentication and selectable encryption to a public specification. This is where the standard advice is simply right. Budget the licence per CPU before you price the project.

If you have many vendors on the floor, you want one API — and there are two ways to get one. OPC UA gives you one by making every device speak the same protocol. The other way is to put the single API in your own application, and let it speak whatever each device speaks best. Which of those you want is a real decision with real trade-offs, and it gets a chapter of its own further down rather than a bullet here.

Before you replace a slow controller, check its request cost and its request size. A six-fold difference per request and a 240-byte limit explained every slow row on my oldest device. Both are cheaper to work around in software than in hardware.

Check the communication cap on every controller you are comparing. It is worth a factor of two and the default is not the same on every family.

And measure on your oldest, busiest machine, not your newest. Watch the cycle time rather than the response time. A controller that is answering you at the cost of its own scan will not tell you so; it will just get less precise, and you will look for the cause somewhere else.

And if security is the axis you care about most, it may be worth waiting a few weeks. On 13 October 2026 we are publishing a series on S7 security, written with Siemens under coordinated disclosure. This article has deliberately stayed on what these protocols cost. That one is about what they actually protect, and if you are close to a decision that turns on the difference, read it before you commit.

13. Choosing per device, without rewriting per device

A disclosure first, because this is published by a company with an interest in its conclusions. I work on ToddySoft Connect, which provides drivers for every protocol here, and every number came out of its test suite. The tables came first and the argument came out of them. You should still know about my interest before you weigh my conclusions.

OPC UA’s strongest selling point is not speed and it is not really security either. It is the unified API. One client, one address space, one set of semantics, across controllers from vendors who agree on nothing else. That is a genuine engineering good and the free protocols offer nothing like it.

But look at how it delivers that. The abstraction lives on the wire. You get one API because every device speaks the same protocol, which means every device must speak it, including the ones that are worst at it. You cannot buy the unification without paying for the transport on every controller you own.

There is another place to put that abstraction. Move it up a layer, so you have one API in your application and many protocols underneath, and the protocol becomes a deployment detail instead of an architectural commitment.

So the new cell runs OPC UA, signed and encrypted, because there the security costs almost nothing. The twelve-year-old controller at the far end of the hall runs S7CommPlus, still named, still free. The old S7-300 that speaks nothing else runs S7Comm. Your application does not know the difference.

That is what ToddySoft Connect does, and it is the reason this bench exists at all: every protocol in this article is a driver we maintain, so the cost of each one is a question we have to keep answering anyway.

Read the next section before you take that as a recommendation.

13.1. What that does not buy you

A uniform API is not uniform capability. No driver layer conjures a type a device does not have. What it can do is turn a rewrite into a constraint you design around.

An address-based protocol has a cost no time column shows. The same value sits at a different address on every controller, so your configuration changes per device, and a stale address does not fail loudly. It reads the neighboring bytes and decodes them into a plausible number.

And you do not get OPC UA’s information model. A unified access API gives you names, types and values. It does not give you semantics: what this machine is, what this value means, in a form another vendor’s software can interpret. If you need that, you need OPC UA.

13.2. What the bench has already changed

Look again at what the tables say about S7Comm. It is the cheapest thing a SIMATIC can do, and its one crippling limitation is that it cannot name anything.

For a long time I read that as a statement about what S7Comm cannot do. It is not. It is a statement about where the layout comes from — and there is a controller on the other end of the wire that already knows it.

So an S7 driver is now being built that fetches names and types over S7CommPlus, then switches to S7Comm to move the data. The expensive protocol once, at connect, for the thing it is uniquely good at. The cheap one for every read after that. It would also keep the merging, because the reads are positional again.

Whether it lands anywhere near that is exactly the kind of claim this article exists to be sceptical of. It will get the same treatment as everything else: measured, with the off-switch built in first.

14. What I got wrong

Every section above has a correction folded into it. Not one was caught by review. Each was caught by somebody asking an obvious question of a number that looked fine.

"Merging is worth 7.6x on S7Comm, and 18x on Modbus." Both figures were diluted by the benchmark’s own design. Each run read every value individually before reading them in batches, and nothing can merge a request carrying one value — so both halves of every comparison paid that phase, which capped every ratio near the number of batched rounds. Measured batched read against batched read, it is 100 to 130 times on Modbus and 14 to 30 on S7Comm.

"Our optimizer makes some controllers slower." JVM warm-up, not the driver. The benchmark was measuring its own compilation.

"The old S7-1212C is slow because it is old." Its 240-byte request size and its six-fold per-request cost explain it. The PLC program contributes nothing measurable.

"Unmerged Modbus cannot read 2000 values." It can. That was our own five-second timeout, applied once per read request.

"OPC UA hits a request-size limit somewhere above a hundred values." No. One controller’s server needs 12.4 ms per value and ran past the deadline we set it.

"OPC UA is a hundred times slower than S7Comm on a SIMATIC." Only for the shape we first read it in. Every OPC UA value was an array element, which is that protocol’s worst case. Read as separate variables, the G2’s server was about eight times faster, and level with S7CommPlus.

Six mistakes, one shape. Five of the six were our own configuration mistaken for a property of somebody else’s product. Which is worth remembering the next time a benchmark tells you something convenient.

15. What this does not tell you

  • The load test is one bench, with an assumed cap. Every run assumed a 20% communication cap, the S7-1200 default, rather than reading it from each CPU. The controllers were running an idle program. A controller already at ninety percent of its cycle has less room than any figure here suggests, which is the case that matters most and the one I cannot measure for you.

  • Reads only — but not as big a gap as that sounds. Every figure here is a read. Separate driver tests put writing at about the same speed as reading on almost every protocol. There is one asymmetry, and it follows from everything above: when merging reads, a client can over-read and throw away the bytes between your values. When merging writes it cannot, because those bytes belong to something else and writing them would destroy it. So where your values have gaps between them, a write splits into more chunks than the same read does, and gains less from merging.

  • One bench, one vendor, five controllers. The ratios travel better than the absolute times.

  • Contiguous values, so the tables are a bracket rather than a prediction. Every figure reads one array in order. That makes the merged rows the best a client could possibly do, and the unmerged rows the worst. Your plant sits somewhere between the two, and where it sits depends on how your data block is laid out. The gap between the two rows is the size of the prize, not the size of your winnings. For OPC UA it runs the other way: the array is its worst case.

  • The ratio ceiling is partly the benchmark. Twenty batched rounds caps it near 21; the shape is real, the height is not all protocol.

  • Every figure states its communication cap, and you should too. It is worth a factor of two.

16. The point

There is no best protocol. There is a controller with a budget, a network with a budget, a threat model, and a set of options that spend those budgets very differently.

But the biggest number in this whole article is not a protocol choice at all. It is the difference between a client that asks well and one that does not — a hundredfold on Modbus, and up to thirty on S7Comm, on the same wire, to the same controller, for the same data.

Group your reads. Hold your connections open. Measure before you commit. And when somebody tells you a protocol is the answer, ask them what it costs.

Back to Blog

Related Posts

View All Posts »
Why Pay for the OPC UA Server on a Beckhoff Controller?

Why Pay for the OPC UA Server on a Beckhoff Controller?

Every Beckhoff controller already speaks ADS. OPC UA is a licensed server on top. Measured from the controller’s own figures: reading natively by address barely registered, up to 2500 values every 10 ms. A busy OPC UA server put its work in the PLC task. Resolving names in the runtime on every read cost up to 46% of its CPU. A look at what the OPC UA license buys, and where each way of reading puts its work.

Live chat is disabled. Cookie consent is required to use this feature. If you don't see a consent banner, an ad-blocker or privacy extension may be preventing it from appearing. Live-Chat ist deaktiviert. Cookie-Zustimmung ist erforderlich, um diese Funktion zu nutzen. Falls Sie kein Zustimmungsbanner sehen, verhindert möglicherweise ein Werbeblocker oder eine Datenschutz-Erweiterung dessen Anzeige.