Every server operator faces the same tension: you need logs to debug problems, detect intrusions, and understand what your infrastructure is doing. But those same logs — if left unconstrained — become a detailed record of every person who interacted with your system, what they did, when they did it, and from where. That is not an operations tool. That is a surveillance archive.
Most operators never think about this. They deploy default logging configurations, accumulate months of detailed access logs, and never consider what they are building until a data breach exposes those logs, a legal request demands them, or a privacy regulation audit asks why they are storing visitor IP addresses for three years when their retention policy says 30 days.
If you chose privacy hosting in Switzerland because you take data protection seriously, your logging infrastructure should reflect that choice. Hosting on a Swiss VPS or dedicated Swiss server puts your data under the Swiss Federal Act on Data Protection (FADP) — one of the strongest privacy frameworks in the world. But that jurisdictional protection only matters if your operational practices match the standard. Storing detailed visitor logs indefinitely on Swiss hardware does not make you privacy-compliant. It makes you a privacy-focused host running a surveillance operation.
This guide covers how to build logging and monitoring infrastructure that gives you everything you need for operations — debugging, security detection, performance analysis — without retaining the data that turns your servers into a liability.
Let us look at what a standard server setup actually collects. A fresh Linux installation with Nginx, SSH, and a typical web application generates logs that contain:
• Nginx access logs: Full client IP address, timestamp, requested URL (including query parameters), HTTP status code, response size, referrer header, user agent string — for every single request
• Nginx error logs: Client IPs, failed request details, upstream errors that may include application-level data
• SSH auth logs: Source IP addresses, usernames attempted, authentication methods, timestamps of every connection attempt
• Systemd journal: Service startup/shutdown times, configuration details, process arguments that may include secrets passed on the command line
• Application logs: User session identifiers, request payloads, database query parameters, error stack traces that include variable contents, API keys accidentally logged in debug mode
• Mail logs: Sender/recipient addresses, source IPs, message IDs, delivery timestamps
• Kernel logs: Network interface events, firewall drops with source/destination IPs, hardware events
Now multiply this by time. A moderately busy website generating 50,000 requests per day produces approximately 500 MB of access logs per month. After a year, you have 6 GB of detailed records mapping IP addresses to browsing behaviour, timestamps, and content interactions. Under GDPR and the FADP, IP addresses are personal data. You have been building a personal data repository without a legal basis, a retention policy, or a data processing agreement.
The default combined log format in Nginx looks like this:
# Default Nginx combined log format
# This is what most servers are running right now
log_format combined '$remote_addr - $remote_user [$time_local] '
'"$request" $status $body_bytes_sent '
'"$http_referer" "$http_user_agent"';
# Actual log entry:
# 203.0.113.47 - - [28/Sep/2026:10:15:23 +0000] "GET /account/settings?user=john.doe HTTP/2.0" 200 4523 "https://www.example.com/dashboard" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"
That single line contains: a real IP address (personal data), a timestamp (behaviour tracking), a URL with a username in the query string (PII leakage), the previous page visited (browsing behaviour), and a device fingerprint via the user agent string. Your web server writes one of these every time someone loads a page. That is not logging. That is profiling.
The first step toward privacy-first logging is asking a question that most operators never consider: what will you actually use this data for?
Most access log data is never read. It sits on disk until logrotate compresses and eventually deletes it. The small fraction that gets examined falls into three categories:
• Debugging: Why did this request return a 500? What was the request URL, what happened upstream?
• Security: Is someone brute-forcing SSH? Is there an unusual pattern of requests suggesting reconnaissance?
• Performance: Which endpoints are slow? What is the request volume trend?
Notice what is not on this list: identifying specific visitors by IP address. For the vast majority of server operators, knowing that a request came from 203.0.113.47 is not useful. Knowing that a request to /api/checkout returned a 502 and took 12 seconds — that is useful. The IP address is a habit, not a requirement.
Here is an Nginx log format that provides operational value without personal data:
# Privacy-first Nginx log format
# Retains operational data, removes identifying information
# Option 1: Anonymised IP (truncate last octet for IPv4, last 80 bits for IPv6)
map $remote_addr $anonymised_addr {
~(?P<ip>\d+\.\d+\.\d+)\.\d+ $ip.0;
~(?P<ip>[^:]+:[^:]+:[^:]+): $ip::;
default 0.0.0.0;
}
# Option 2: No IP at all — just a request hash for deduplication
map $request_id $req_hash {
default $request_id;
}
log_format privacy_ops '$anonymised_addr - [$time_local] '
'"$request_method $uri" $status $body_bytes_sent '
'$request_time $upstream_response_time '
'"$server_name"';
# What this logs:
# 203.0.113.0 - [28/Sep/2026:10:15:23 +0000] "GET /account/settings" 200 4523 0.045 0.032 "www.example.com"
#
# What changed:
# - IP anonymised (last octet zeroed)
# - Query string removed from URI ($uri instead of $request)
# - Referrer removed (browsing behaviour tracking)
# - User agent removed (device fingerprinting)
# - Request time and upstream time ADDED (actually useful for debugging)
# - Server name ADDED (useful for multi-domain setups)
access_log /var/log/nginx/access.log privacy_ops;
# For specific locations where you need NO logging at all
location /health {
access_log off;
}
location /api/heartbeat {
access_log off;
}
This format gives you everything you need for debugging (HTTP method, URI path, status code, response time, upstream time) and removes everything that constitutes personal data (full IP, user agent, referrer, query parameters). The anonymised IP preserves enough network information to identify regional issues (is the problem affecting a specific /24 subnet?) without identifying individuals.
IP anonymisation is more nuanced than most implementations suggest. There are several approaches, each with different privacy guarantees and operational trade-offs:
Truncation (zeroing the last octet for IPv4) is the simplest and most widely used. It reduces a /32 address to a /24 network, preserving geographic and network-level information while removing individual identification. This is what Google Analytics uses and what most European DPAs have accepted as adequate anonymisation for web analytics.
# IPv4 truncation with Nginx map
map $remote_addr $ip_anon {
~(?P<ip>\d+\.\d+\.\d+)\.\d+ $ip.0;
default 0.0.0.0;
}
# For stronger anonymisation, truncate two octets (/16)
map $remote_addr $ip_anon_strong {
~(?P<ip>\d+\.\d+)\.\d+\.\d+ $ip.0.0;
default 0.0.0.0;
}
Hashing (SHA-256 of the IP, optionally with a daily-rotating salt) preserves the ability to correlate requests from the same source within a time window without revealing the actual IP. This is useful for abuse detection — you can see that "hash X made 10,000 requests in 5 minutes" without knowing who hash X is.
#!/bin/bash
# /usr/local/bin/log-anonymise.sh
# Pipe Nginx logs through this for hash-based IP anonymisation
# Usage in Nginx: access_log syslog:server=unix:/var/run/log-anon.sock privacy_raw;
# Daily salt rotation — prevents long-term correlation
SALT=$(date +%Y-%m-%d)
SECRET=$(cat /etc/nginx/anonymisation-secret)
while IFS= read -r line; do
# Extract the IP (first field)
IP=$(echo "$line" | awk '{print $1}')
# Generate HMAC — keyed hash prevents rainbow table attacks
HASHED=$(echo -n "${IP}${SALT}${SECRET}" | sha256sum | cut -c1-16)
# Replace IP with hash prefix
echo "$line" | sed "s/^[^ ]*/anon_${HASHED}/"
done
Complete removal — replacing the IP with a static placeholder — provides maximum privacy but eliminates network-level debugging capability. For applications where you genuinely do not need any source identification (static content sites, public documentation), this is the cleanest approach.
# Complete IP removal in Nginx
log_format no_ip '- - [$time_local] "$request_method $uri" $status '
'$body_bytes_sent $request_time';
# Every request logs as:
# - - [28/Sep/2026:10:15:23 +0000] "GET /docs/api" 200 8432 0.012
The right choice depends on your threat model. If you are running offshore hosting for clients who specifically chose Swiss jurisdiction for privacy protection, complete removal or hashing with a short-lived salt is appropriate. If you need to maintain some abuse-detection capability on a Swiss VPS, truncation to /24 provides a reasonable balance.
Web server logs are the easy part. Application logs are where the real privacy risks hide — because developers log whatever is convenient for debugging, and "convenient" often means "the entire request object including the user's personal data."
A typical application error log might contain:
# What your application actually logs during an error:
ERROR 2026-09-28 10:22:15 payment.service - Payment processing failed
Request: POST /api/v2/payments
User: john.doe@example.com (user_id: 48291)
IP: 203.0.113.47
Payload: {"card_number": "4532-XXXX-XXXX-1234", "amount": 99.50,
"currency": "CHF", "billing_address": "Bahnhofstrasse 42, 8001 Zürich"}
Error: StripeError: card_declined - insufficient_funds
Stack trace:
at PaymentProcessor.charge (payment.js:142)
at OrderController.processPayment (order.js:89)
...
That log entry contains: an email address, a user ID, an IP address, a partial card number, a financial transaction amount, a physical address, and the information that this specific person's card was declined for insufficient funds. If this log is breached, it is not just a security incident — it is a financial privacy violation.
The solution is a structured sanitisation layer between your application and your log storage:
# Python: Privacy-aware logging configuration
import logging
import re
import hashlib
from datetime import date
class PrivacySanitiser(logging.Filter):
"""
Logging filter that strips PII from log messages before they reach
any handler (file, syslog, log aggregator).
"""
# Patterns to detect and sanitise
PATTERNS = [
# Email addresses
(re.compile(r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}'),
'[EMAIL_REDACTED]'),
# IPv4 addresses (anonymise to /24)
(re.compile(r'\b(\d{1,3}\.\d{1,3}\.\d{1,3})\.\d{1,3}\b'),
r'\1.0'),
# Credit card numbers (any format)
(re.compile(r'\b\d{4}[-\s]?\d{4}[-\s]?\d{4}[-\s]?\d{4}\b'),
'[CARD_REDACTED]'),
# Swiss phone numbers
(re.compile(r'\+?41[\s.-]?\d{2}[\s.-]?\d{3}[\s.-]?\d{2}[\s.-]?\d{2}'),
'[PHONE_REDACTED]'),
# International phone numbers (broad pattern)
(re.compile(r'\+\d{1,3}[\s.-]?\d{2,4}[\s.-]?\d{3,4}[\s.-]?\d{3,4}'),
'[PHONE_REDACTED]'),
# Swiss postal addresses (basic pattern)
(re.compile(r'\b\d{4}\s+[A-ZÄÖÜa-zäöü]+(?:\s+[A-ZÄÖÜa-zäöü]+)*\b'),
'[ADDRESS_REDACTED]'),
# IBAN numbers
(re.compile(r'\b[A-Z]{2}\d{2}[\s]?\d{4}[\s]?\d{4}[\s]?\d{4}[\s]?\d{4}[\s]?\d{0,2}\b'),
'[IBAN_REDACTED]'),
# JWT tokens (prevent session hijacking from logs)
(re.compile(r'eyJ[A-Za-z0-9_-]+\.eyJ[A-Za-z0-9_-]+\.[A-Za-z0-9_-]+'),
'[JWT_REDACTED]'),
# Bearer tokens
(re.compile(r'Bearer\s+[A-Za-z0-9_-]+'),
'Bearer [TOKEN_REDACTED]'),
# API keys (common patterns)
(re.compile(r'(?:api[_-]?key|apikey|token|secret)["\s:=]+["\']?[A-Za-z0-9_-]{20,}',
re.IGNORECASE),
'[APIKEY_REDACTED]'),
]
def filter(self, record):
"""Sanitise the log message before it reaches any handler."""
if isinstance(record.msg, str):
for pattern, replacement in self.PATTERNS:
record.msg = pattern.sub(replacement, record.msg)
# Also sanitise args if they contain strings
if record.args:
sanitised_args = []
for arg in (record.args if isinstance(record.args, tuple)
else (record.args,)):
if isinstance(arg, str):
for pattern, replacement in self.PATTERNS:
arg = pattern.sub(replacement, arg)
sanitised_args.append(arg)
record.args = tuple(sanitised_args)
return True
class PrivacyAwareFormatter(logging.Formatter):
"""
Formatter that adds request context without PII.
Replaces user identifiers with pseudonymised hashes.
"""
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
self._daily_salt = None
self._salt_date = None
@property
def daily_salt(self):
today = date.today()
if self._salt_date != today:
self._daily_salt = hashlib.sha256(
f"log-salt-{today.isoformat()}".encode()
).hexdigest()[:16]
self._salt_date = today
return self._daily_salt
def pseudonymise_user_id(self, user_id):
"""Create a daily-rotating pseudonym for user IDs in logs."""
if user_id is None:
return "anonymous"
h = hashlib.sha256(
f"{user_id}-{self.daily_salt}".encode()
).hexdigest()[:12]
return f"user_{h}"
# Configure privacy-aware logging
def setup_privacy_logging(app_name: str, log_file: str):
logger = logging.getLogger(app_name)
logger.setLevel(logging.INFO)
# Add privacy filter to ALL handlers
privacy_filter = PrivacySanitiser()
# File handler with rotation
handler = logging.handlers.RotatingFileHandler(
log_file,
maxBytes=50 * 1024 * 1024, # 50 MB
backupCount=3, # Keep only 3 rotated files
)
handler.addFilter(privacy_filter)
handler.setFormatter(PrivacyAwareFormatter(
'%(asctime)s %(levelname)s %(name)s - %(message)s'
))
logger.addHandler(handler)
return logger
# Usage:
log = setup_privacy_logging("payment-service", "/var/log/app/payment.log")
# Developer writes this (bad habit, but the filter catches it):
log.error(f"Payment failed for user john@example.com from 203.0.113.47")
# What actually gets written to disk:
# 2026-09-28 10:22:15 ERROR payment-service - Payment failed for user [EMAIL_REDACTED] from 203.0.113.0
The sanitisation filter acts as a safety net. Developers will log PII — it is inevitable when debugging at 2 AM. The filter ensures that even when they do, the personal data is stripped before it reaches the log file. Defence in depth: the best practice is to not log PII in the first place, but the filter catches what slips through.
On modern Linux systems, systemd's journal captures vastly more data than most administrators realise. By default, it stores persistent logs with no size limit until the disk fills up, and it captures every stdout/stderr line from every service, every kernel message, and every audit event.
# Check how much journal data you are storing
journalctl --disk-usage
# Output: Archived and active journals take up 4.8G on disk.
# See what kinds of data are in your journal
journalctl --output=json-pretty | head -50
# Reveals: _HOSTNAME, _MACHINE_ID, _BOOT_ID, _TRANSPORT,
# _SYSTEMD_UNIT, MESSAGE, _PID, _UID, _GID, _COMM, _EXE,
# _CMDLINE (full command line including arguments!)
The _CMDLINE field is particularly dangerous. If anyone starts a process with credentials on the command line (a depressingly common practice), the journal captures and stores those credentials permanently:
# This is bad practice, but it happens constantly:
mysql -u admin -pMySecretPassword123 production_db
# The journal now contains:
# _CMDLINE=mysql -u admin -pMySecretPassword123 production_db
# Stored on disk, potentially forever, readable by root
Here is how to configure the journal for privacy-first operation:
# /etc/systemd/journald.conf
# Privacy-hardened journal configuration
[Journal]
# Limit storage — do not accumulate years of data
SystemMaxUse=500M
SystemKeepFree=2G
SystemMaxFileSize=50M
# Maximum retention — logs older than this are deleted
MaxRetentionSec=7d
# Rate limiting — prevents log flooding from capturing excessive data
RateLimitIntervalSec=30s
RateLimitBurst=1000
# Compress stored journals
Compress=yes
# Do not forward to syslog (avoid double-storage)
ForwardToSyslog=no
# Seal journal entries (detect tampering)
Seal=yes
# Apply the configuration
sudo systemctl restart systemd-journald
# Verify the new limits are active
journalctl --disk-usage
# Manually vacuum old entries right now
sudo journalctl --vacuum-time=7d
sudo journalctl --vacuum-size=500M
Seven days of journal retention is sufficient for most operational debugging. If you need longer retention for security monitoring, route specific security-relevant events to a separate, access-controlled log store — do not keep the entire system journal for months because you might need one authentication failure event from six weeks ago.
Under both GDPR and the Swiss FADP, you must have a documented retention policy for any data that contains personal information — and logs containing IP addresses, user identifiers, or session tokens qualify. "We keep logs forever because disk is cheap" is not a retention policy. It is a compliance violation waiting for an audit.
A practical retention schedule for privacy-first operations:
• Anonymised access logs (no PII): 90 days — sufficient for trend analysis, performance baselines, and capacity planning
• Application error logs (sanitised): 30 days — most bugs are investigated within hours; a month provides margin for intermittent issues
• Authentication logs (security): 30 days with full detail, then 90 days anonymised — balance security investigation needs with privacy
• Firewall/IDS logs: 30 days detailed, 180 days summarised (connection counts per source /24, not individual IPs)
• System journal: 7 days — operational debugging only
Implementing this with logrotate:
# /etc/logrotate.d/privacy-nginx
/var/log/nginx/access.log {
daily
rotate 90
compress
delaycompress
missingok
notifempty
dateext
dateformat -%Y-%m-%d
sharedscripts
postrotate
[ -f /var/run/nginx.pid ] && kill -USR1 $(cat /var/run/nginx.pid)
endscript
}
/var/log/nginx/error.log {
daily
rotate 30
compress
delaycompress
missingok
notifempty
dateext
dateformat -%Y-%m-%d
sharedscripts
postrotate
[ -f /var/run/nginx.pid ] && kill -USR1 $(cat /var/run/nginx.pid)
endscript
}
# /etc/logrotate.d/privacy-auth
/var/log/auth.log {
daily
rotate 30
compress
delaycompress
missingok
notifempty
dateext
dateformat -%Y-%m-%d
# After 30 days, a separate cron job anonymises and moves to cold storage
}
# /etc/logrotate.d/privacy-app
/var/log/app/*.log {
daily
rotate 30
compress
delaycompress
missingok
notifempty
dateext
dateformat -%Y-%m-%d
copytruncate
}
Crucially: enforce deletion, do not just configure rotation. Logrotate compresses old files, but verify that files beyond your retention window are actually being removed. A misconfigured rotate directive or a missing maxage setting can result in years of accumulated logs that violate your own retention policy.
#!/bin/bash
# /usr/local/bin/log-retention-audit.sh
# Weekly cron job to verify log retention compliance
# Alerts if any log files exceed their maximum retention
MAX_ACCESS_DAYS=90
MAX_APP_DAYS=30
MAX_AUTH_DAYS=30
MAX_JOURNAL_DAYS=7
violations=0
# Check Nginx access logs
old_access=$(find /var/log/nginx/ -name "access.log*" -mtime +${MAX_ACCESS_DAYS} | wc -l)
if [ "$old_access" -gt 0 ]; then
echo "VIOLATION: ${old_access} access log files older than ${MAX_ACCESS_DAYS} days"
violations=$((violations + old_access))
fi
# Check application logs
old_app=$(find /var/log/app/ -name "*.log*" -mtime +${MAX_APP_DAYS} 2>/dev/null | wc -l)
if [ "$old_app" -gt 0 ]; then
echo "VIOLATION: ${old_app} app log files older than ${MAX_APP_DAYS} days"
violations=$((violations + old_app))
fi
# Check auth logs
old_auth=$(find /var/log/ -name "auth.log*" -mtime +${MAX_AUTH_DAYS} | wc -l)
if [ "$old_auth" -gt 0 ]; then
echo "VIOLATION: ${old_auth} auth log files older than ${MAX_AUTH_DAYS} days"
violations=$((violations + old_auth))
fi
# Check journal size
journal_bytes=$(journalctl --disk-usage 2>&1 | grep -oP '\d+\.\d+[MG]')
echo "Current journal size: ${journal_bytes}"
if [ "$violations" -gt 0 ]; then
echo "TOTAL VIOLATIONS: ${violations} — cleaning up"
# Auto-cleanup with verification
find /var/log/nginx/ -name "access.log*" -mtime +${MAX_ACCESS_DAYS} -delete
find /var/log/app/ -name "*.log*" -mtime +${MAX_APP_DAYS} -delete 2>/dev/null
find /var/log/ -name "auth.log*" -mtime +${MAX_AUTH_DAYS} -delete
sudo journalctl --vacuum-time=${MAX_JOURNAL_DAYS}d
echo "Cleanup complete"
else
echo "All log retention policies compliant"
fi
Logging is one half of observability. Monitoring — collecting metrics about your infrastructure's performance and health — is the other half. The good news: most monitoring metrics are inherently privacy-neutral. CPU usage, memory consumption, disk I/O, network throughput, request counts, error rates — none of these contain personal data in their raw form.
The risk comes from labels and dimensions. A metric like http_requests_total is harmless. A metric like http_requests_total{client_ip="203.0.113.47", path="/api/user/profile", user_id="48291"} is a surveillance tool wearing a monitoring hat.
# Prometheus: Privacy-safe metric labels
# GOOD — operational labels only
http_request_duration_seconds_bucket{method="GET", path="/api/payments", status="200", le="0.5"}
http_request_duration_seconds_bucket{method="POST", path="/api/orders", status="201", le="1.0"}
# BAD — PII in labels (creates high cardinality AND privacy violation)
# http_request_duration_seconds{client_ip="203.0.113.47", user="john.doe"}
# Rule: Never use IP addresses, user IDs, email addresses, or session tokens
# as Prometheus label values. Besides the privacy issue, high-cardinality
# labels will also destroy your Prometheus performance.
For Nginx, export metrics using the stub_status module or the VTS (Virtual Host Traffic Status) module, which provide aggregate counters without per-request detail:
# Nginx VTS module configuration for privacy-safe metrics
# Provides request counts, response times, and traffic volume
# WITHOUT logging individual requests
http {
# Enable VTS module
vhost_traffic_status_zone;
# Expose metrics endpoint (restrict to monitoring network)
server {
listen 127.0.0.1:9113;
location /metrics {
vhost_traffic_status_display;
vhost_traffic_status_display_format prometheus;
# Only accessible from localhost (Prometheus scrapes locally)
allow 127.0.0.1;
deny all;
}
}
# Filter out client-identifying dimensions
vhost_traffic_status_filter_by_set_key $server_name server_zone;
# Do NOT add: vhost_traffic_status_filter_by_set_key $remote_addr client
}
For application-level metrics, the same principle applies. Instrument your code with counters, histograms, and gauges that describe system behaviour, not user behaviour:
# Python: Privacy-safe application metrics with Prometheus client
from prometheus_client import Counter, Histogram, Gauge
# GOOD: Operational metrics
request_duration = Histogram(
'app_request_duration_seconds',
'Request duration in seconds',
['method', 'endpoint', 'status_code'], # No user labels
buckets=[0.01, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5, 5.0, 10.0]
)
payment_total = Counter(
'app_payments_total',
'Total payment attempts',
['status', 'currency'], # Payment outcome, not who paid
)
active_sessions = Gauge(
'app_active_sessions',
'Number of active user sessions',
[], # Total count only, no per-user tracking
)
db_pool_usage = Gauge(
'app_db_pool_connections',
'Database connection pool usage',
['pool', 'state'], # active/idle/waiting
)
# BAD: PII in metrics — NEVER do this
# user_actions = Counter('user_actions_total', 'Actions per user',
# ['user_id', 'action', 'ip_address'])
The hardest balance in privacy-first logging is security monitoring. Intrusion detection, brute-force protection, and anomaly detection all traditionally rely on detailed logging of network activity — exactly the kind of data that creates privacy risks. The question is whether you can detect threats without building a comprehensive record of legitimate activity.
The answer is yes, but it requires a different architecture. Instead of logging everything and searching later, you process events in real-time, extract security signals, and discard the raw data:
#!/bin/bash
# /usr/local/bin/auth-monitor.sh
# Real-time SSH auth monitoring with privacy preservation
# Detects brute force without storing individual connection records
# Temporary state file — reset daily
STATE_FILE="/tmp/auth-monitor-state"
ALERT_THRESHOLD=10 # Alert after 10 failed attempts from same /24
# Process auth log in real-time
tail -F /var/log/auth.log | while read -r line; do
# Only process failed authentication
if echo "$line" | grep -q "Failed password\|Invalid user\|authentication failure"; then
# Extract and anonymise source IP to /24
FULL_IP=$(echo "$line" | grep -oP '\d+\.\d+\.\d+\.\d+' | head -1)
if [ -n "$FULL_IP" ]; then
ANON_IP=$(echo "$FULL_IP" | sed 's/\.[0-9]*$/.0/')
TIMESTAMP=$(date +%s)
# Count failures from this /24 in the last hour
echo "${TIMESTAMP} ${ANON_IP}" >> "$STATE_FILE"
# Clean entries older than 1 hour
HOUR_AGO=$((TIMESTAMP - 3600))
if [ -f "$STATE_FILE" ]; then
awk -v cutoff="$HOUR_AGO" '$1 > cutoff' "$STATE_FILE" > "${STATE_FILE}.tmp"
mv "${STATE_FILE}.tmp" "$STATE_FILE"
fi
# Count recent failures from this network
COUNT=$(grep -c "$ANON_IP" "$STATE_FILE" 2>/dev/null || echo 0)
if [ "$COUNT" -ge "$ALERT_THRESHOLD" ]; then
# Alert with anonymised data only
logger -p auth.warning "SECURITY: ${COUNT} failed auth attempts from ${ANON_IP}/24 in last hour"
# Optional: auto-block the /24 with fail2ban or iptables
# iptables -A INPUT -s ${ANON_IP}/24 -p tcp --dport 22 -j DROP
fi
fi
fi
done
For more sophisticated intrusion detection, tools like OSSEC and Wazuh can be configured to process events in real-time, fire alerts based on rules, and discard the raw event data after analysis:
# /var/ossec/etc/ossec.conf (excerpt)
# OSSEC/Wazuh configuration for privacy-first security monitoring
<ossec_config>
<!-- Log analysis with privacy controls -->
<localfile>
<log_format>syslog</log_format>
<location>/var/log/auth.log</location>
</localfile>
<!-- Alert storage settings -->
<alerts>
<log_alert_level>3</log_alert_level>
<!-- Only store alerts, not raw events -->
</alerts>
<!-- Archive settings — DISABLE raw event archiving -->
<global>
<logall>no</logall>
<logall_json>no</logall_json>
<!-- Do not archive all received events -->
<!-- Only fired alerts are stored -->
</global>
<!-- File integrity monitoring (no PII involved) -->
<syscheck>
<frequency>7200</frequency>
<directories check_all="yes">/etc,/usr/bin,/usr/sbin</directories>
<directories check_all="yes">/var/www/html</directories>
<ignore>/etc/mtab</ignore>
<ignore>/etc/resolv.conf</ignore>
</syscheck>
</ossec_config>
The key configuration: logall=no tells OSSEC/Wazuh to process all events through its rules engine but only store the events that trigger an alert. Legitimate activity is analysed and immediately discarded. Only suspicious activity is retained for investigation. This is the difference between "we monitor for threats" and "we record everything in case we need it later."
If you operate multiple servers — a web frontend, an application backend, a database server, perhaps a caching layer — you probably want centralised logging. Aggregating logs into a single searchable system (ELK stack, Loki, Graylog) makes debugging cross-service issues dramatically easier.
But centralised logging amplifies every privacy problem. Instead of PII scattered across individual servers (each requiring separate access), you now have a single database containing personal data from every system, searchable with a query language, accessible to anyone with dashboard credentials.
Privacy-first log aggregation requires:
# Rsyslog: Anonymise before forwarding to central log server
# /etc/rsyslog.d/90-privacy-forward.conf
# Load string manipulation module
module(load="mmexternal")
# Template that anonymises IPs before forwarding
template(name="PrivacyForward" type="list") {
property(name="timestamp" dateFormat="rfc3339")
constant(value=" ")
property(name="hostname")
constant(value=" ")
property(name="syslogtag")
# Message with IP anonymisation applied
property(name="msg" regex.expression="([0-9]+\\.[0-9]+\\.[0-9]+)\\.[0-9]+"
regex.submatch="1" regex.nomatchmode="FIELD")
constant(value=".0 ")
property(name="msg")
constant(value="\n")
}
# Forward to central log server with anonymisation
# Use TLS for transport security
action(
type="omfwd"
target="logserver.internal"
port="6514"
protocol="tcp"
template="PrivacyForward"
StreamDriver="gtls"
StreamDriverMode="1"
StreamDriverAuthMode="x509/name"
StreamDriverPermittedPeers="logserver.internal"
)
# CRITICAL: Use TLS for log transport
# Unencrypted syslog over UDP is a data breach waiting to happen
# Every log message crosses the network in plaintext
If you use a log aggregation platform like Grafana Loki, configure it to enforce privacy at the query level:
# Loki configuration for privacy-first log aggregation
# /etc/loki/config.yml
auth_enabled: true # MANDATORY — no anonymous access to logs
server:
http_listen_port: 3100
ingester:
lifecycler:
ring:
kvstore:
store: inmemory
replication_factor: 1
chunk_idle_period: 5m
chunk_retain_period: 30s
schema_config:
configs:
- from: 2026-01-01
store: boltdb-shipper
object_store: filesystem
schema: v11
index:
prefix: index_
period: 24h
storage_config:
boltdb_shipper:
active_index_directory: /var/lib/loki/boltdb-shipper-active
cache_location: /var/lib/loki/boltdb-shipper-cache
shared_store: filesystem
filesystem:
directory: /var/lib/loki/chunks
# Retention — enforce at the storage level
# Logs are automatically deleted after the retention period
table_manager:
retention_deletes_enabled: true
retention_period: 720h # 30 days maximum
limits_config:
# Prevent high-cardinality labels that could track users
max_label_name_length: 1024
max_label_value_length: 2048
max_label_names_per_series: 15
# Rate limits to prevent log flooding
ingestion_rate_mb: 10
ingestion_burst_size_mb: 20
# Query limits
max_query_length: 720h # Cannot query beyond retention window
Fail2ban is a staple of server security — it monitors log files for patterns indicating brute-force attacks and temporarily bans the offending IP addresses. By necessity, it processes and temporarily stores IP addresses. Here is how to configure it for minimum data retention:
# /etc/fail2ban/jail.local
# Privacy-hardened Fail2ban configuration
[DEFAULT]
# Short ban time — block the attack, then forget the attacker
bantime = 1h
findtime = 10m
maxretry = 5
# Use firewall-based blocking (no persistent storage of banned IPs)
banaction = iptables-multiport
# Database settings — minimise stored data
# By default, Fail2ban stores ban history in SQLite
# Reduce retention to minimum needed for repeat-offender detection
dbpurgeage = 86400 # Purge ban records after 24 hours (default is 86400*7)
[sshd]
enabled = true
port = ssh
filter = sshd
logpath = /var/log/auth.log
maxretry = 3
bantime = 2h
# Use incremental banning for repeat offenders
# without long-term storage of IP addresses
bantime.increment = true
bantime.maxtime = 24h
bantime.factor = 2
[nginx-http-auth]
enabled = true
port = http,https
filter = nginx-http-auth
logpath = /var/log/nginx/error.log
maxretry = 5
bantime = 1h
# Verify Fail2ban database retention
# Check current database size and purge old records
sudo fail2ban-client get dbpurgeage
# Output: 86400
# Manually purge old records now
sudo fail2ban-client set dbpurgeage 86400
# Check what is currently stored
sudo fail2ban-client status sshd
# Shows current ban count, not historical data
Running privacy-first logging infrastructure on a Swiss dedicated server provides advantages that go beyond technical configuration:
Legal framework alignment. The Swiss FADP requires that data processing — including logging — have a legal basis and be proportionate to its purpose. Operating your logs under Swiss law means your minimisation efforts are not just best practice; they are meeting a specific legal standard that courts and regulators have interpreted consistently. The FADP's data minimisation principle (Article 6(2)) directly supports the technical choices described in this guide: collect only what is necessary, retain only as long as needed, protect what you store.
Jurisdictional protection for log data. Even anonymised or minimised logs, when combined with other data sources, can potentially re-identify individuals. Swiss jurisdiction under the FADP means that any foreign request for your log data — even the minimised version — must go through the Swiss mutual legal assistance process, providing judicial review and proportionality assessment before disclosure. On offshore hosting platforms in less protective jurisdictions, your carefully minimised logs could be bulk-subpoenaed without meaningful judicial oversight.
No compelled backdoor logging. Some jurisdictions require or pressure hosting providers to implement enhanced logging beyond what operators configure — data retention mandates, lawful intercept capabilities, or metadata collection requirements. Switzerland's FADP does not impose blanket data retention obligations on hosting providers. Your logging configuration is your logging configuration. If you configure Nginx to anonymise IPs and retain logs for 30 days, that is what happens. No regulatory framework compels you to secretly maintain a detailed copy.
Physical security for log infrastructure. If your centralised log server is on dedicated hardware in a Swiss data centre, your log data benefits from the same physical security controls (biometric access, 24/7 monitoring, locked cabinets) and the same jurisdictional protection as your primary infrastructure. Contrast this with cloud-based log aggregation services where your log data — including any residual PII that survives your sanitisation filters — is stored on shared infrastructure in a jurisdiction you may not control.
Before implementing changes, understand what your servers are currently collecting. Run this audit on every server:
#!/bin/bash
# /usr/local/bin/privacy-log-audit.sh
# Audit current logging configuration for privacy risks
# Run as root on each server
echo "=== Privacy Log Audit: $(hostname) ==="
echo "Date: $(date -Iseconds)"
echo ""
# 1. Total log volume
echo "--- Log Storage ---"
du -sh /var/log/ 2>/dev/null
echo ""
# 2. Oldest log files (retention check)
echo "--- Oldest Log Files ---"
find /var/log/ -type f -name "*.log*" -printf '%T+ %p\n' 2>/dev/null | sort | head -10
echo ""
# 3. Nginx log format (PII check)
echo "--- Nginx Log Format ---"
grep -r "log_format" /etc/nginx/ 2>/dev/null
echo ""
# 4. Check for full IPs in recent access logs
echo "--- IP Addresses in Access Logs ---"
if [ -f /var/log/nginx/access.log ]; then
echo "Unique IPs in last 1000 lines:"
tail -1000 /var/log/nginx/access.log | awk '{print $1}' | sort -u | wc -l
echo "Sample (first 5):"
tail -1000 /var/log/nginx/access.log | awk '{print $1}' | sort -u | head -5
fi
echo ""
# 5. Journal retention settings
echo "--- Journal Configuration ---"
grep -v "^#\|^$" /etc/systemd/journald.conf 2>/dev/null
echo "Journal size:"
journalctl --disk-usage 2>&1
echo ""
# 6. Logrotate configuration
echo "--- Logrotate Retention ---"
for conf in /etc/logrotate.d/*; do
echo " $conf:"
grep -E "rotate |maxage " "$conf" 2>/dev/null | head -2
done
echo ""
# 7. Check for sensitive data in recent logs
echo "--- Potential PII in Logs (sampling) ---"
echo "Email patterns:"
grep -rchP '[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}' /var/log/nginx/ /var/log/app/ 2>/dev/null | awk '{sum+=$1} END {print sum " matches"}'
echo "API key patterns:"
grep -rchP 'api[_-]?key|bearer|token|secret' /var/log/app/ 2>/dev/null | awk '{sum+=$1} END {print sum " matches"}'
echo ""
# 8. Check for command-line secrets in journal
echo "--- Command-line Secrets in Journal ---"
journalctl --no-pager -q --since "24 hours ago" | grep -ciP 'password|secret|token|api.?key' 2>/dev/null
echo "potential occurrences in last 24h"
echo ""
echo "=== Audit Complete ==="
Run this audit. The results will probably surprise you. Most operators find that they are storing more personal data in their logs than in their databases — and with far less access control.
Privacy-first logging requires deliberate engineering effort. Default configurations in every major web server, application framework, and Linux distribution are designed for maximum operational visibility, not for data minimisation. Changing these defaults means accepting some trade-offs:
You will occasionally wish you had the full IP address when debugging a network issue. You will sometimes want the user agent string to investigate a browser-specific bug. You will miss the convenience of grepping through months of detailed access logs when investigating an incident that started six weeks ago.
These are real operational costs. The question is whether they outweigh the risks of maintaining a comprehensive surveillance archive — the breach exposure, the compliance liability, the legal discovery risk, and the fundamental inconsistency of hosting on privacy-focused Swiss infrastructure while logging like you are running an advertising network.
For operators who chose a Swiss VPS or high-bandwidth Swiss server specifically for privacy — for themselves or their users — the answer should be clear. Your logging infrastructure should reflect the same values that led you to Swiss hosting in the first place. Collect what you need, protect what you store, and delete what you no longer require.
That is not just good privacy practice. It is good engineering. The best log is one that tells you exactly what broke, without telling you anything about who was there when it happened.