At the tutorial of CIGRE study committee B5, three renowned international experts in protection and digital substations showed the next step in the industry's evolution. Protection functions are gradually separating from dedicated terminals and moving onto server computing platforms. But the most interesting discussion today is no longer about whether this is technically possible, but about how to design, test and operate such systems — and who is ultimately responsible for their operation.

In the hall of the Paris Palais des Congrès, at the panel table sit Alex Apostolov, editor-in-chief of PAC World magazine and a specialist at OMICRON electronics; Ratan Das, a specialist at Siemens Energy and convener of the IEEE working group that produced the IEEE C37.300 guide on centralised protection and control; and David MacDonald, solutions and standardisation architect at GE Grid Automation and convener of CIGRE working group B5.84 on virtualised protection devices.

The B5 study committee tutorial was devoted to the development of the architectures of protection, automation and control systems, with special attention to the process bus and the virtualisation of protection devices.

The three talks lined up almost into a single historical sequence.

First, the link of protection devices with the primary equipment becomes digital thanks to the process bus. Then the protection functions separate from the numerous physical terminals and are centralised. Finally, the next step — virtualisation: the protection function turns into a software application, and the server into a computing platform that can later be replaced without changing the protection algorithm itself.

Over some twenty-odd years, the industry has travelled the path from the first demonstrations of transmitting digitised instantaneous values over IEC 61850 to virtualised protection systems running on server platforms.

And the nature of the questions has changed together with the technology.

Today the main question is no longer «will it work?», but: how do you own it, how do you test and maintain it — and who is responsible for the system if its components are supplied by different manufacturers?

Twenty years from a demonstration to industrial use

Alex Apostolov's first talk was devoted to the process bus.

At its core is the CIGRE technical brochure TB 949, «Process Bus in Protection, Automation & Control Systems», which summarises the experience the industry has accumulated in applying the process bus: architectural solutions, redundancy, real projects, the results of surveys of utilities and manufacturers, economic aspects and further directions of development.

The story turned out to be longer than is sometimes imagined.

The work that later led to IEC 61850 began back in the 1990s. In 2004, the 9-2LE profile became an important practical milestone in the development of transmitting digitised instantaneous values and in interoperability between equipment from different manufacturers. Apostolov recalled how, at the CIGRE session of that same year, ABB, Siemens and Omicron were already demonstrating devices working together with SV streams.

However, about a decade passed between the demonstration of the technology and its wide industrial use.

The industry looked closely at the new architecture, accumulated experience, launched pilot projects. From the mid-2010s the number of deployments began to grow rapidly.

«We are at a turning point: almost all the world's major utilities have set a course for digitalisation».

Behind this conclusion lies a rather large body of material: more than 200 publications analysed, around 60 projects reviewed and the results of an industry survey.

Most of the projects studied relate not to the construction of new facilities but to the refurbishment of existing substations. This matters: the process bus is gradually ceasing to be a technology only for new, «ideal» digital facilities and is entering the practice of modernising existing substations.

B5 CIGRE tutorial: the panel and a slide on the evolution of the process bus
Fig. 1. The B5 CIGRE study committee tutorial: at the panel table — Alex Apostolov, David MacDonald and Ratan Das. On screen — the stages of process-bus development (1994 → 2004 → 2015+); Alex Apostolov presenting.

A short glossary

For the discussion that follows, it is useful to define a few terms that are fixed in international terminology.

  • MU / SAMU — devices that convert analogue current and voltage signals into digital streams of instantaneous SV values.
  • SCU — a unit interfacing with the switchgear that provides binary I/O.
  • PIU — a process-interface unit that combines the functions of converting analogue and binary signals.

As Apostolov noted, the terminology here is still used not entirely consistently.

  • PRP and HSR — network-redundancy mechanisms with seamless recovery of data transmission.
  • PTP — the IEEE 1588 high-precision time-synchronisation protocol, which makes it possible to transmit the time scale directly over the Ethernet network and, in many architectures, to do without separate pulse-synchronisation lines.
The «key terms» slide from Alex Apostolov’s talk
Fig. 2. The glossary from Alex Apostolov’s talk: MU/SAMU, SCU, LPIT, SV/GOOSE, PRP/HSR, PTP.

What drives the process bus — and what holds it back

The reasons for adoption can be divided into economic and technological.

The economics — less copper cabling and installation work, more compact cabinets and rooms, a reduced scope of work during refurbishment and a potential reduction in maintenance and upgrade costs.

The technological advantages — improved personnel safety, architectural flexibility, scalability, wider possibilities for monitoring the state of equipment and a basis for further functional integration.

But the barriers are quite concrete too.

Utilities lack their own experience of working with the technology. The difficulties of getting equipment from different manufacturers to work together are especially noticeable at the design and configuration stage. Cybersecurity places additional demands on the architecture. New methods of diagnostics and testing are needed.

And, perhaps, one of the most important questions is tools.

The IEC 61850 standard itself makes it possible to build many different architectures, yet the complexity of configuration is still too often shifted directly onto the engineer. Apostolov separately stressed the need for tools that should hide the technological complexity of IEC 61850 from the user, rather than forcing them to work with it manually all the time.

«The market needs tools that will hide the technological complexity of IEC 61850 from the user, rather than forcing them to work with it manually all the time».

A separate issue is the economics.

A digital substation is often assessed by comparing the cost of individual components of a traditional and a digital solution. But the real benefit only shows up when the total cost of ownership is assessed: design, cabling, construction, commissioning, maintenance, refurbishment and the subsequent replacement of equipment.

And it is precisely the life cycle that becomes especially interesting at the next stage of development.

From ten thousand connections to centralised protection

Ratan Das began his talk with a personal story.

In 1981, having only just joined a utility, he was given the task of preparing the wiring diagram between devices at a facility with three 500 MW generating units.

Around 10,000 connections.

The designer drew up the diagrams, then they were checked, and then all of this had to be physically installed and tested on site.

The next generations of protection gradually reduced this complexity.

Electromechanical relays gave way to static ones, then microprocessor terminals appeared. IEC 61850 made it possible to replace a significant part of the wired links with digital exchange.

The next logical question arose about a decade ago: if the analogue measurements and binary signals are already available over the process bus, does every bay necessarily have to have its own computing terminal?

Thus the concept of Centralized Protection and Control — CPC — appeared.

The interface units stay near the primary equipment. They receive measurements, states and commands. And the protection and control algorithms are moved onto a centralised computing platform at station level.

In Das's words, this can be regarded as the next generation in the development of protection systems.

«When in 2013 we wrote a paper on a server-based protection system, many looked at it as a wild fantasy. Today at the stands of this session, protection runs precisely on servers».

But centralisation solves only part of the task.

The functions are already assembled on a common computing platform, yet the software may still remain tightly bound to the manufacturer's specific hardware platform.

The next step is to separate one from the other.

Ratan Das on centralised protection and functionality independent of hardware
Fig. 3. Ratan Das (Siemens Energy): the layered software architecture of protection, automation and control systems with functionality independent of hardware (FIH) — application, platform and driver layers.

Functions apart, hardware platform apart

This is precisely the idea of functionality independent of hardware — FIH.

The basic principle is simple: replacing the computing platform should not require a substantial rework of the applied protection functions.

Of course, there is no such thing as absolute independence. A new platform may require drivers, a compatibility check, configuration of the runtime environment.

But the protection algorithm itself should not have to be created anew each time merely because the processor manufacturer has discontinued an old generation or the server hardware has changed.

This is what constitutes the fundamental change in the life cycle.

Today a protection terminal is effectively a single product: the protection algorithms, the processor, the operating system, the network interfaces, the power supply and the enclosure are supplied together.

When the hardware becomes obsolete, the whole terminal often has to be replaced.

In a virtualised system, the applied function and the computing platform begin to live different life cycles.

For the power industry this is especially important. Primary equipment is operated for decades, whereas computing platforms develop much faster.

As a result, several generations of protection computing hardware may change over the service life of the primary equipment.

With the new architecture, changing the protection function itself each time is no longer mandatory.

There may be more testing, and less work on site

At first glance it seems that virtualisation should reduce the amount of testing.

Das made an important qualification: it is more correct to put it quite differently.

You can even carry out more testing — but move a significant part of it from the site to the laboratory.

A software application can be comprehensively tested independently of a specific facility.

If subsequently only the computing platform changes, while the code of the protection function itself remains the same, the full scope of functional testing of the application no longer necessarily has to be repeated directly at the substation.

The bulk of the verification of a new platform can be done in advance.

This potentially seriously changes the very process of replacing equipment.

Today, replacing a terminal often means a project, approvals, taking equipment out of service, installation, configuration and a full scope of checks.

In a software-defined protection system, replacing the hardware platform may in the future become a much more routine operation.

But along with this, a fundamental organisational question arises: who gives the warranty for the whole system and how is responsibility distributed among the participants?

The discussion returned to it after the talks.

What exactly virtualisation gives

The third talk — by David MacDonald, convener of CIGRE working group B5.84 on virtualised protection devices — was devoted directly to virtualisation.

The main difference of a virtualised system from an ordinary centralised one lies not only in the use of a server.

Between the applied functions and the hardware platform there appears a virtualisation layer, which makes it possible to separate the software applications from the specific server hardware and to isolate them from one another.

For virtual machines, the basic element of such a layer is the hypervisor — a piece of software that makes it possible to run several isolated virtual machines simultaneously on one physical server.

The applied protection functions can run in separate virtual machines. Another approach is containerisation, where several applications use a common operating system but are isolated from one another by software means. A combined variant is also possible — for example, placing containerised applications inside a virtual machine.

For systems where the use of applications from different manufacturers is envisaged, virtual machines are especially interesting thanks to the stricter separation of software environments.

A separate virtual protection function can be deployed, tested or updated independently of other functions running on the same server.

It is precisely this capability that creates the preconditions for systems in which applications from different manufacturers can run on a common computing platform.

In a traditional centralised system there is significantly less such freedom: the functions are combined, but the whole solution remains a single hardware-software system from one supplier.

The hardware platform does not disappear

Here it is important not to fall into a terminological trap.

Independence from the hardware platform does not mean that any server will do for protection.

The requirements for the computing platform remain extremely serious.

You have to take into account the performance of the processor and memory, the network interfaces, the redundancy of power and disks, the operating conditions, electromagnetic compatibility, the means of remote management, the redundancy scheme and the reliability requirements.

In addition, for each virtual device you have to define the required computing resources and the characteristics of the virtual network.

That is, virtualisation does not do away with the engineering design of the hardware.

It changes something else: the hardware platform ceases to determine the applied functionality of the protection itself.

This is an important distinction.

David MacDonald on application isolation in a virtualised protection system
Fig. 4. David MacDonald (GE Grid Automation, convener of CIGRE WG B5.84): application isolation in a vPAC — containers, virtual machines and the hybrid option.

Two redundancy models

For a server-based protection system, different approaches to redundancy can be used.

The most obvious — two independent servers on which identical instances of the protections run simultaneously. Each is capable of issuing its own trip command.

A more complex variant — a cluster of several servers.

In this case a group of nodes is managed as a single computing system. On the failure of a server, its applications can be automatically restarted on another node.

To make decisions about the state of the cluster's members, an odd number of nodes is used — for example, three or five. This makes it possible to determine a majority and exclude the failed node.

It is also possible to migrate a virtual machine from one physical server to another.

And here one of the fundamental features of the new architecture arises: what is now made redundant is not necessarily the physical device itself, but the ability to execute the function.

For a traditional protection system these two concepts practically coincided.

For a virtualised one — no longer.

Microseconds: where virtualisation meets real time

However, a server-based protection system places demands on the computing infrastructure that differ greatly from ordinary information systems.

It is not enough to ensure a short function-execution time on average.

Protection requires a predictable reaction time and a limited spread of delays even under adverse conditions.

Ordinary operating systems are first of all optimised for the efficient use of resources, not for strict temporal determinism.

Virtualisation itself can also introduce additional delays: a network packet passes through several software layers before reaching the protection application.

To shorten this path and reduce delays, special mechanisms of direct and accelerated access to the network interfaces are used — in particular, PCI passthrough, SR-IOV and DPDK.

A detailed analysis of these technologies is the subject of a separate article. What matters here is something else: in a virtualised protection system, the network path and the way data is delivered to the application become just as much a part of the engineering design as the choice of processor, the allocation of computing resources or the tuning of real-time operation.

But MacDonald's main conclusion was not about the choice of a specific technology.

Determinism cannot simply be declared. It has to be measured and confirmed by testing.

How the open SEAPATH platform is built

As a vivid example of an architecture, MacDonald examined the open LF Energy SEAPATH platform.

Here the virtual protection devices sit on top of a set of standard tools for virtualisation and management of the computing infrastructure.

Linux with a kernel for real-time operation, KVM, QEMU and libvirt are used. For the virtual network, the corresponding software tools are applied, and for building a fault-tolerant infrastructure — clustering mechanisms.

The allocation of computing resources plays an important role.

Processor cores can be pinned to particular tasks. The NUMA architecture is taken into account — the organisation of processors' access to the main memory in multiprocessor server platforms. Thread priorities and interrupt handling are managed.

That is, the server ceases to be simply «a computer on which protection has been launched».

It becomes a specially tuned computing environment for executing critical functions in real time.

Testing one protection without stopping the others

Virtualisation offers very interesting possibilities together with the standard IEC 61850 Test and Simulation mechanisms.

Imagine that the virtual devices of several bays are running on one server.

One of them needs to be tested.

It can be switched to test mode and fed simulated SV and GOOSE streams from a test set.

Meanwhile, the neighbouring virtual devices continue to work with the facility's real data.

The test GOOSE messages are marked accordingly, so the actuating devices that remain in normal mode should not execute the commands contained in them.

Thus it becomes possible to localise the testing of an individual function inside an operating system.

For a traditional centralised architecture, such independence of individual software components is achieved with significantly more difficulty.

Determinism has to be proven under load

The working group also considers special testing of the computing platform.

It is measured how much the actual task-execution time differs from the expected.

Such tests are performed many times and under an artificially created load on the processor, memory and network subsystem.

The most telling is the end-to-end test.

The test set generates a stimulus, the virtual protection processes it and returns a GOOSE trip command.

The full path of the signal is measured — from the stimulus to the responding command.

Moreover, the computing platform is deliberately placed under adverse conditions of competition for resources.

It is precisely here that the fundamental difference between an ordinary server application and a protection function becomes visible.

For protection, what matters is not the best result and not even the average one.

What matters is the guaranteed characteristic under the worst foreseen conditions.

Cybersecurity: there are no benefits without new risks

Centralisation and virtualisation simultaneously create new means of protection and new threats.

On the one hand, the isolation of virtual machines makes it possible to limit the spread of failures and the consequences of the compromise of individual applications.

Trusted boot, integrity checking and signing of software images, rights segregation, multi-factor authentication and encryption can be applied.

On the other hand, new critical components appear: the virtualisation environment itself, the node's operating system, the centralised-management tools, the software supply chain.

In addition, the consequences of a successful attack on a common server infrastructure can potentially be much larger in scale than the compromise of an individual terminal.

Therefore, to claim that a virtualised system is «safer» in itself would be too simple.

It is more correct to say that virtualisation changes the threat model and the set of protective means.

Questions from the floor: from reliability to responsibility

The most interesting part of the tutorial began after the microphones were opened.

The questions from the floor showed well how much the nature of the professional discussion has changed.

The argument was no longer about the possibility of virtualisation as such, but about its practical limits.

«Aren't we creating a new single point of failure?»

The first natural question concerned reliability.

In a traditional digital substation, dozens of physical devices are distributed across the bays. The failure of a specific terminal is usually localised to a limited set of functions.

In a virtualised system, a significant part of the functions may depend on just a few servers.

Doesn't a new single point of failure arise?

The speakers' answer was that the computing infrastructure can be made redundant significantly more deeply and more economically than a large number of dedicated devices.

A protection function can have several simultaneously running instances placed on different servers. On the failure of one node, operation can continue on another.

But the utility representatives rightly drew attention to the other side, too.

The failure of a common component is indeed potentially capable of affecting many functions at once.

Therefore you cannot simply compare «two terminals» and «two servers».

A reliability analysis of the whole architecture is needed: the servers, the switches, the network cards, the power supplies, the virtualisation environment, the management tools and the applications themselves.

Virtualisation does not do away with traditional reliability engineering.

It only provides other ways of building redundancy.

«And who is responsible if the software and the server come from different manufacturers?»

Perhaps the most practical question came from a representative of Scottish and Southern Electricity Networks (SSEN, UK) — a UK electricity network operator.

Suppose the virtual protection was developed by manufacturer A.

The server platform was supplied by manufacturer B.

The virtualisation environment — by manufacturer C.

The protection operated incorrectly.

Whom should the customer call?

This is one of the key unsolved tasks of the new architecture.

The speakers noted that a certain analogue already exists today in digital substations with equipment from different manufacturers.

Devices from different companies are combined by a common IEC 61850 network, but responsibility for the system as a whole must still be assigned to a specific party — the customer or a designated system integrator.

For a virtualised protection system this problem only becomes more noticeable.

Consequently, the industry needs not only technical standards for interfaces.

It needs rules for the qualification of computing platforms, for application compatibility, for diagnostics, for testing and for the contractual distribution of responsibility.

In effect, this is about forming a new operating model for protection systems.

It is better to speak not of independence but of freedom from constraints

An interesting formulation was proposed by a participant in the discussion from the Electricity Supply Board (ESB, Ireland) — an Irish energy company.

In his view, the expression «independence from the hardware platform» does not quite accurately describe the problem.

A traditional protection function does not so much «depend» on the hardware platform as it is constrained by it.

The service life of protection algorithms is often determined not by their becoming morally obsolete at all.

Power supplies, voltage converters, processor modules fail. The manufacturer discontinues a specific hardware platform — and the whole device has to be changed along with it.

Even if the algorithms themselves fully satisfy the utility.

Virtualisation potentially breaks this link.

The function remains, and the computing platform can be updated independently.

In this sense it is indeed rather a matter not of the complete absence of hardware dependence, but of freeing the applied functionality from the constraints of a specific generation of equipment.

Where independence from the hardware platform ends

A representative of Schweitzer Engineering Laboratories (SEL, USA) — a well-known manufacturer of devices and solutions for protection, automation and control — raised an even more interesting technical question.

Standard protection functions can be virtualised and executed on a standard server platform.

But there are methods in which the algorithm is tightly bound to dedicated measuring equipment.

For example, travelling-wave protections require an extremely high sampling rate.

In such cases, transmitting a huge stream of data to a central server may be irrational.

One option is to perform such processing directly in the process-interface unit.

But here, too, not everything is so simple.

If the function uses data from several remote points, a local computation on one device no longer solves the task.

In the end the discussion led to a more important conclusion than the idea of «moving everything to the server»: virtualisation does not mean the mandatory centralisation of absolutely all functions.

Different algorithms can be executed at different levels.

Functions that use exclusively local data and require minimal delay can stay near the process.

Functions that use information from several bays — be executed at substation level.

System functions — at an even higher level.

That is, the future may turn out to be not fully centralised, but rather a hierarchically distributed software-defined protection system in which the place of execution of each function is chosen based on its requirements.

And what about medium voltage?

A separate question concerned distribution networks.

At high voltage levels, a protection terminal is a comparatively expensive element. At medium voltage, the cost of individual devices is substantially lower.

Does the economic sense of virtualisation hold there?

The speakers believe that it does.

Modern switchgear assemblies increasingly acquire built-in digital measurement and control means. In the future, the process-interface unit may become part of the switching apparatus itself.

And the protection functions are then added in software.

The result is another level of separation of life cycles: the procurement of primary equipment is separated from the procurement and subsequent development of protection and automation functions.

It is precisely mass distribution substations that may in the future become one of the largest areas of application for software-defined protection systems.

The bottom line

The most telling thing at the tutorial was not spoken from the slides.

Throughout the whole block of open discussion, almost no one asked: «Why is virtualised protection needed at all?»

The questions were different.

How do you ensure the required determinism? How do you make the server infrastructure redundant? How do you test an individual function without affecting the others? What do you leave right at the process, and what do you move to the server? How does cybersecurity change? And, finally, who is responsible for the result if the application, the virtualisation environment and the hardware platform are supplied by different companies?

For a technology that only about ten years ago many perceived as a rather bold concept, this is perhaps the best indicator of maturity.

Protection really is beginning to leave the «iron».

But the main challenge of the next stage is no longer to prove the fundamental possibility of virtualisation.

The industry will have to learn to design, qualify, test and operate protection as a software-defined system.

And judging by what the specialists argued about at CIGRE 2026, this transition has already begun.