From Manufacturer Cloud Services to Controlled Infrastructure
Modern IP cameras are much more than digital eyes. Behind the lens is a small computer with a processor, operating system and network connection – sometimes even with its own AI capabilities. This creates an apparent contradiction: a camera designed to protect a building can itself become a digital attack surface. Limiting this risk depends not only on the camera, but on the architecture of the entire video surveillance system.
When choosing a surveillance camera, other questions usually come first: How high is the resolution? How good is the image at night? How large is the field of view? Does the camera provide motion detection?
From an IT security perspective, however, another question needs to be asked:
What is the camera actually allowed to do on our network?
This question is becoming increasingly important. Modern video surveillance and conventional computer security can no longer be treated as entirely separate disciplines.
A Modern Camera Is a Computer
A modern IP camera typically has its own processor, memory, network interface and operating system. It also runs software and network services for various tasks.
These may include:
- video transmission,
- a web-based administration interface,
- user accounts,
- time synchronisation,
- motion detection,
- firmware updates,
- smartphone applications,
- remote access and
- connections to cloud services.
From an IT security perspective, such a camera is therefore a networked computer.
And like any other computer, a camera can contain security vulnerabilities.
Outdated firmware, weak passwords, unnecessary network services or insecure remote access can turn a device intended to provide security into a security risk itself.
The issue is not limited to whether an unauthorised person might be able to view a camera feed.
An equally important question is:
What happens if someone gains control of the camera?
Could the camera then reach other devices on the corporate network? Could it communicate freely with servers on the Internet? Could it attack other cameras?
This brings us to one of the fundamental questions of IT security: What do we actually trust?
Security Is a Question of Trust
It would be too simplistic to say that cloud systems are inherently insecure while self-hosted systems are inherently secure.
A professionally operated cloud service can be very well protected. Conversely, a poorly maintained private server can present significant security risks.
The fundamental difference therefore lies elsewhere:
Who controls the infrastructure, and whom does the operator have to trust?
Depending on the architecture of the video surveillance system, the answer can be very different.
Model 1: The Convenient Manufacturer Cloud
Many modern cameras are designed to make installation as simple as possible.
The camera is connected to the network, an app is installed and an account is created with the manufacturer. Within a short time, the user can access the camera from almost anywhere.
In simplified form, the architecture looks like this:
Camera
│
local network
│
Internet
│
manufacturer cloud
│
Internet
│
app / user
This is convenient and can be entirely appropriate for many applications.
The convenience does, however, have a consequence: the chain of trust becomes longer.
The operator relies, among other things, on the manufacturer, its servers, its account management, its update procedures and the long-term availability of the service.
This does not automatically make the solution insecure.
It does mean that part of the security control lies outside the operator’s own infrastructure.
Model 2: Keeping the Recordings In-House
One alternative is to store the video data on a system under the operator’s own control.
A NVR – Network Video Recorder – is commonly used for this purpose. Put simply, an NVR is the central video recorder of an IP camera system.
The cameras transmit their video streams to this server over the network:
Camera ───── network ───── NVR
│
▼
local storage
The continuous video stream therefore does not have to be transmitted to an external cloud service for recording.
An NVR can be a dedicated appliance, but it can also run on a self-hosted Linux server.
Open-source solutions such as Frigate, Shinobi or ZoneMinder make this type of self-hosted operation possible.
This gives the operator greater control.
At the same time, however, the operator becomes responsible for updates, user accounts, backups and securing the server.
More control therefore also means more responsibility.
Model 3: Open Software on the Camera
Open camera firmware such as OpenIPC takes this concept one step further.
OpenIPC is based on Linux and supports various processor platforms used in IP cameras.
Its key advantage is not that open-source software is automatically secure.
It is not.
Open software can contain bugs and security vulnerabilities just like proprietary software.
The difference lies primarily in transparency and configurability.
With open software, it is possible in principle to examine which software is being used. Functions can be adapted, unnecessary components removed and additional security mechanisms integrated.
This is particularly interesting for small, purpose-built Linux systems: they do not necessarily need every feature a manufacturer may have included to serve as many different customers and use cases as possible.
The system can instead be configured more specifically for its intended purpose.
Open source is therefore not a promise of security. It does, however, provide possibilities for transparency and control that are often unavailable with completely closed software.
Even Linux Does Not Mean Complete Control
There is, however, an important limitation.
A camera consists of more than its visible operating system.
Below Linux, additional components may be operating, such as firmware for individual hardware components, bootloaders, device drivers, image processors or other parts of the underlying chip.
Not all of these components are necessarily open to inspection.
On conventional PCs, this issue is familiar from additional system components such as the Intel Management Engine. A typical IP camera does not necessarily contain a directly comparable system.
The underlying problem nevertheless remains:
Even when a camera runs an open Linux system, not every hardware component is automatically open and fully controllable.
This leads to an interesting conclusion.
Perhaps we do not need to trust the camera completely in the first place.
Model 4: Giving the Camera Its Own Network Zone
Instead of assuming that a camera will never be compromised, the network can be designed so that a compromised camera has as little reach as possible.
For this purpose, the cameras are separated from the normal corporate network.
In simplified form:
Camera 1 ─┐
Camera 2 ─┼── dedicated camera network
Camera 3 ─┘
│
firewall
│
▼
NVR
A firewall controls which connections are permitted to leave the camera network.
The office network, file servers and other corporate systems remain inaccessible to the cameras.
The principle is straightforward:
A camera receives only the network connections it actually needs to perform its task.
For example:
Camera → NVR ALLOW
Camera → local time server ALLOW
Camera → Internet BLOCK
Camera → office network BLOCK
Camera → server network BLOCK
This does not automatically make the camera secure.
It does, however, limit the consequences of a possible compromise.
Why Should Cameras Be Able to Talk to Each Other?
With multiple cameras, another interesting question arises.
Does Camera 1 need to communicate directly with Camera 2?
In many installations, the answer is no.
Suitable network switches can therefore be configured so that individual cameras can reach the NVR but cannot communicate directly with one another.
Camera 1 ──X── Camera 2
Camera 1 ──X── Camera 3
Camera 2 ──X── Camera 3
Camera 1 ─────► NVR
Camera 2 ─────► NVR
Camera 3 ─────► NVR
If one camera is compromised, this makes it more difficult for an attacker to move directly to another camera.
This illustrates a fundamental security principle: a device should receive only the permissions and connections it genuinely requires.
Model 5: Moving the Security Boundary Outside the Camera
A Linux-based camera can run its own firewall. A secure VPN such as WireGuard can also run directly on suitable camera firmware.
That can be useful.
It has, however, an inherent weakness.
If an attacker gains complete control of the camera, they may also be able to modify its firewall, routing or VPN configuration.
The device being protected would then also control part of its own security boundary.
For systems with higher security requirements, another approach can therefore be used.
A separate security gateway is placed between the camera and the rest of the network.
This could, for example, be a small, hardened Linux computer with two network interfaces.
Camera
│
│
▼
┌────────────────────┐
│ Security Gateway │
│ │
│ Firewall │
│ Network rules │
│ VPN │
│ Logging │
└─────────┬──────────┘
│
▼
external network
The critical security boundary is now outside the camera.
This is an important distinction.
Even if the camera changes its own network configuration, it does not automatically gain control of the separate security gateway.
Its network connection leads only to that system.
What Happens in the Worst Case?
Let us deliberately assume an unfavourable scenario:
The camera has been completely compromised.
On a conventional network, it might now attempt to reach other devices or external servers.
With the separated security model, however, it encounters an independent boundary:
POTENTIALLY UNTRUSTED
Camera
│
▼
════════ SECURITY BOUNDARY ════════
Security Gateway
/ Firewall
│
▼
permitted systems
The camera can attempt to contact an arbitrary Internet address.
The gateway drops the connection.
It can attempt to reach an office computer.
The gateway drops the connection.
It can disable its own firewall.
The external firewall remains unaffected.
This is the core principle of this security model:
The camera does not decide whom it is allowed to communicate with. The network does.
Model 6: Connecting Remote Sites Through a Secure Tunnel
This principle becomes particularly interesting when cameras are installed at remote sites.
The cameras might be located at a customer’s premises while the NVR or monitoring centre is operated elsewhere.
In this case, the security gateway can establish an encrypted WireGuard VPN tunnel.
CUSTOMER SITE
Camera 1 ─┐
Camera 2 ─┼── isolated network
Camera 3 ─┘
│
▼
Security Gateway
│
WireGuard VPN
│
▼
════════════ INTERNET ════════════
│
▼
own NVR
own monitoring centre
The cameras themselves do not require unrestricted Internet access.
Only the gateway communicates externally.
And even within the VPN, not every connection needs to be permitted automatically. The gateway can continue to define which systems and services are accessible.
If the VPN tunnel fails, the system can also be designed according to a fail-closed principle.
This means that the failure of the protected connection does not automatically result in an unprotected fallback connection over the open Internet.
Model 7: Keeping Video Analysis Local as Well
Operating a private NVR creates another possibility: artificial intelligence can also run within the operator’s own infrastructure.
One example is Frigate.
Frigate combines video recording with local object detection. Depending on the hardware, video streams can automatically be analysed for particular objects.
These may include:
- people,
- vehicles,
- animals or
- movement within defined areas.
The important point is:
The video material does not inherently have to be transmitted to an external AI provider for this analysis.
Processing can take place on local hardware.
Camera
│
▼
Security Gateway
│
▼
own NVR
│
▼
local AI
│
▼
event / alert
Both recording and substantial parts of automated analysis can therefore remain within the operator’s own infrastructure.
Detecting a Person Is Not the Same as Detecting a Security Incident
The capabilities of AI should nevertheless be viewed realistically.
An AI system may be able to determine:
“There is a person in this image.”
That does not automatically mean:
“A burglary is taking place.”
Video analysis becomes more interesting when several pieces of information are combined.
For example:
person detected
+
restricted area
+
02:37 a.m.
+
no scheduled activity
↓
possible security event
Local AI can therefore help filter large volumes of video data and highlight unusual events.
Determining whether an actual threat exists and deciding what response is appropriate remain separate tasks.
In professional security services, this division of labour can be particularly useful: technology detects and filters – people assess and decide.
Time and Updates Are Part of Security Too
An isolated camera network does not mean that cameras should operate without supporting infrastructure.
A seemingly minor example is accurate timekeeping.
If Camera 1 records an incident at 02:37 while Camera 2 records the same event at 02:42, reconstructing the sequence later becomes unnecessarily difficult.
A local time server can therefore be useful.
Updates must not be forgotten either.
Disconnecting a camera from the Internet and then leaving it without security updates for years would not be a sound security strategy.
Updates can instead be obtained in a controlled manner, verified and then installed within the protected environment.
Isolation is not a substitute for maintenance.
Several Layers of Protection Instead of One Supposedly Perfect Camera
This brings us to the underlying principle of the entire architecture.
No single measure provides complete security.
Open source can provide transparency and configurability, but it does not prevent software vulnerabilities.
A firewall restricts network communication, but it does not prevent the camera itself from being compromised.
WireGuard protects the communication channel, but it cannot make a compromised endpoint trustworthy again.
A self-hosted NVR provides greater control over storage, but the NVR itself must then be secured.
Local AI reduces dependency on external analysis platforms, but it introduces another software system that must be maintained.
For this reason, several independent layers of protection are combined.
In IT security, this principle is known as Defense in Depth.
If one layer fails, another should continue to limit the potential damage.
From a Simple Cloud System to a Dedicated Security Architecture
The different approaches can be compared in simplified form:
MANUFACTURER CLOUD
Camera
↓
manufacturer cloud
↓
app
LOCAL RECORDING
Camera
↓
own NVR
↓
own storage
SEPARATE CAMERA NETWORK
Cameras
↓
isolated network
↓
firewall
↓
own NVR
EXTERNAL SECURITY BOUNDARY
Cameras
↓
isolated network
↓
Security Gateway
↓
own NVR
REMOTE SITE
Cameras
↓
Security Gateway
↓
WireGuard
↓
own infrastructure
LOCAL DATA PROCESSING
Cameras
↓
Security Gateway
↓
own NVR
↓
local AI
↓
own storage
↓
monitoring centre
This is explicitly not a ranking in which only the final architecture can be considered secure.
Not every application requires the same level of technical complexity.
A single camera monitoring a low-risk area has different requirements from the surveillance of an industrial facility, data centre or other highly protected site.
What matters is that the architecture is chosen deliberately.
The Key Question Is Not: Can We Trust the Camera?
Traditional security considerations often focus on selecting the most trustworthy product possible.
That remains important.
For networked security technology, however, that approach alone is no longer sufficient.
A more robust question is:
How can we design video surveillance so that we do not have to trust the camera unconditionally?
A camera can run open-source firmware and still be isolated.
A camera can use proprietary firmware and still be denied unrestricted Internet access.
A camera can be compromised and still remain separated from office and server networks.
And a camera can modify its own network configuration without gaining control over the firewall of a physically separate security gateway.
That is the difference between an individual security feature and a considered security architecture.
The security of a video surveillance system should therefore not depend on every camera and every individual firmware component being completely trustworthy. A sound security concept instead systematically limits what any single device is able to reach.
Modern video surveillance may still begin at the lens – but it no longer ends there.