Technical report 01 / 10 · Series: anomaly detection in IP flows
When the attack fits in 0.01% of the packets: IP flow entropy in the age of encrypted traffic
Published 21 August 2026
Abstract
Denial-of-service attacks above one terabit per second are no longer exceptional, yet most hostile traffic remains too small and too brief to move any volume meter. This report revisits the central principle of deep IP flow inspection, first presented in [1], and places it against the 2026 landscape: transport encrypted by default, botnets assembled from compromised devices, and campaigns that end before a human operator can react. The argument is that entropy computed over flow metadata measures the shape of traffic rather than its content, which is why it survives encryption. The report describes what entropy represents in terms of concentration and dispersion of flow properties, why a network scan consuming 0.01% of the packets is invisible to volumetric detectors, and which limits of the classical network-level model motivated the move to device-level analysis. This is the first of ten parts.
Keywords: network anomaly, entropy, IP flow, encrypted traffic, intrusion detection, network management.
01The median is small. The tail is enormous.
There is a known distortion in how network attacks get discussed. Headlines record the records, and the records are genuinely striking: in the first half of 2026, Cloudflare reported mitigating 935 network-layer attacks above 1 Tbps, a 519% rise between the first and second quarters[10].
The same report carries the number almost nobody quotes. Over that period, 96.62% of network-layer attacks stayed below 500 Mbps and 90.60% ended in under ten minutes[10]. The distribution is aggressively bimodal: a tail that bends a provider's infrastructure, and a central mass that never shows up on a link utilisation graph.
That central mass is the hard problem. A detector calibrated for the tail steps over it without recording anything. A detector calibrated for it floods the operator with alarms once the tail arrives. What a detection system needs, therefore, is not sensitivity but adjustable sensitivity, a point this series returns to in Part 4.
A small median does not mean a harmless event. A port scan brings nothing down on its own, though it is usually the opening move of someone looking for a vulnerability to exploit[12]. The event that is cheap to generate is precisely the one that precedes the expensive one.
02What encryption left visible
In plain terms
Picture a telephone exchange that cannot listen to the calls but logs every one of them: who called whom, at what time, for how many minutes. Without hearing a single word, that log reveals who dialled three hundred different numbers in two minutes. Flow analysis reads traffic the same way.
Deep packet inspection, which examines the transported content, has lost ground irreversibly. With TLS 1.3[8] and QUIC[9], not only the payload but much of the session metadata that used to travel in the clear is now encrypted by default.
There is, however, one layer encryption cannot erase, because it is the condition for the network to function at all. Routers and switches must know where to forward each packet, so they group traffic into flows and export those records through protocols such as IPFIX[7].
Encryption protects what was said. It does not hide who spoke to whom, when, how often, in what volume, and for how long. A flow record carries source and destination addresses, ports, protocol, packet, byte and flow counts, and timestamps. None of that requires reading a single byte of content.
It is worth being honest about the reach of this choice. A system observing only flow metadata does not detect viruses, Trojan horses, buffer overflows or privilege escalation, precisely because that evidence lives in the content[1]. The scope is narrower, although it is a scope that remains available when everything else closes.
03Entropy, in plain terms
In plain terms
Write down the destination port of every connection entering your network for five minutes. If the list is nearly all the same, entropy is low. If every line holds a different number, entropy is high. Anomalies push that number toward one of the two extremes, and the distance to the extreme is the alarm.
Entropy is a measure physicists invented to quantify disorder and information theory adopted to quantify uncertainty[2]. In traffic analysis it takes on a more direct reading: entropy measures how far a traffic property is concentrated or dispersed.
Given a probability distribution p1, p2, …, pn, Shannon entropy is defined as:
HS = −∑ pi log pi
The value is minimal when nearly all probability piles onto a single point and maximal when it spreads evenly across many. Applied to destination ports, the measure separates a healthy web server, where almost everything arrives at port 443, from a target under scan, where tens of thousands of ports appear once each.
04Concentration and dispersion are the event's signature
What makes entropy interesting is not detecting that something changed. It is that the direction of the change, property by property, describes the type of event.
A TCP flood against a server concentrates the destination address, the destination port, the packet count and the flow count, while dispersing the source port, since the attack tool draws a random high port for each connection. A network scan does nearly the opposite, dispersing addresses and counts with tiny packets.
In [1] this reading was formalised as a textual signature in which each position corresponds to a flow property labelled concentrated (C), dispersed (D) or normal (N). The real distributed attack recorded in that study produced TCCDCCCC, the alpha flow event produced TCCCCCDC and the network scan produced IDDDDD.
These are fingerprints in practice. Two anomalies whose disturbances differ by orders of magnitude remain distinguishable by shape rather than by size.
05Why measuring the whole network is not enough
The classical application of entropy to anomaly detection works at the network level[3][4]. A single distribution is computed by aggregating all traffic observed in the interval, and its deviation from habitual behaviour is checked.
The model performs well as long as there is one event at a time. When several occur simultaneously, however, a dominance effect appears: the attack moving gigabytes rewrites the aggregate distribution, and the discreet event vanishes inside it[4].
The network scan figures in [1] show the scale of the problem. The scanning device probed 254 destinations and accounted, on average, for 0.01% of the packets, 0.001% of the bytes and 0.15% of the flows in the interval. The event was in the traffic. It simply was not in the average.
In the same study, when four distinct anomalies were combined into a single dataset, Shannon entropy applied at the network level identified only one of the four[1]. This is not an implementation defect but a direct consequence of aggregating before measuring.
The proposed way out was to change the unit of observation. Instead of one distribution for the network, entropy is computed over the flows exchanged between each pair of devices, which requires representing the network as a directed graph and inspecting every edge. That is deep IP flow inspection, the subject of Part 3.
The second decision was to replace Shannon entropy with Tsallis entropy[6], whose generalisation introduces an entropic parameter q controlling how much high and low probabilities weigh in the result[5]. That parameter is the sensitivity dial mentioned in Section 1, and Part 4 shows what it costs and what it delivers.
06What comes in the next nine parts
The series walks the full path, from the raw flow record to autonomous operation, always testing the original design against the current landscape.
- IP flows as a data source. NetFlow, IPFIX and what is given up along with the payload.
- The network as a graph. Deep flow inspection and moving the unit of observation to the device.
- The entropic parameter q. Tuning the detector between the tail and the median.
- Adaptive thresholds. Why percentiles replaced mean and standard deviation.
- Signatures and rule-based classification. From detection to diagnosing the event type.
- Simultaneous anomalies. The case where the detector must see more than one thing at once.
- Packet sampling. The cost of seeing 1 packet in 2048.
- Locating source and destination. Heuristics for bidirectional flows with no initiator mark.
- From detection to autonomy. Metrics, key performance indicators and the autonomic control loop.
References
- A. A. Amaral, L. de S. Mendes, B. B. Zarpelão, and M. L. Proença Jr., "Deep IP flow inspection to detect beyond network anomalies," Comput. Commun., vol. 98, pp. 80–96, Jan. 2017, doi: 10.1016/j.comcom.2016.12.007.
- C. E. Shannon, "A mathematical theory of communication," Bell Syst. Tech. J., vol. 27, no. 4, pp. 623–656, Oct. 1948.
- A. Lakhina, M. Crovella, and C. Diot, "Mining anomalies using traffic feature distributions," in Proc. ACM SIGCOMM, Philadelphia, PA, USA, 2005, pp. 217–228.
- G. Nychis, V. Sekar, D. G. Andersen, H. Kim, and H. Zhang, "An empirical evaluation of entropy-based traffic anomaly detection," in Proc. 8th ACM SIGCOMM Conf. Internet Meas. (IMC), Vouliagmeni, Greece, 2008, pp. 151–156.
- A. Ziviani, A. T. A. Gomes, M. L. Monsores, and P. S. S. Rodrigues, "Network anomaly detection using nonextensive entropy," IEEE Commun. Lett., vol. 11, no. 12, pp. 1034–1036, Dec. 2007.
- C. Tsallis, "Possible generalization of Boltzmann-Gibbs statistics," J. Stat. Phys., vol. 52, no. 1–2, pp. 479–487, Jul. 1988.
- B. Claise, B. Trammell, and P. Aitken, "Specification of the IP Flow Information Export (IPFIX) protocol for the exchange of flow information," IETF, RFC 7011, Sep. 2013. [Online]. Available: https://www.rfc-editor.org/rfc/rfc7011
- E. Rescorla, "The Transport Layer Security (TLS) protocol version 1.3," IETF, RFC 8446, Aug. 2018. [Online]. Available: https://www.rfc-editor.org/rfc/rfc8446
- J. Iyengar and M. Thomson, "QUIC: a UDP-based multiplexed and secure transport," IETF, RFC 9000, May 2021. [Online]. Available: https://www.rfc-editor.org/rfc/rfc9000
- Cloudflare, "Cloudflare DDoS threat report H1 2026," Cloudflare Blog, Aug. 2026. [Online]. Available: https://blog.cloudflare.com/ddos-threat-report-2026-h1/
- M. Antonakakis et al., "Understanding the Mirai botnet," in Proc. 26th USENIX Security Symp., Vancouver, BC, Canada, 2017, pp. 1093–1110.
- M. H. Bhuyan, D. K. Bhattacharyya, and J. K. Kalita, "Surveying port scans and their detection methodologies," Comput. J., vol. 54, no. 10, pp. 1565–1581, Oct. 2011.
- A. Sperotto, G. Schaffrath, R. Sadre, C. Morariu, A. Pras, and B. Stiller, "An overview of IP flow-based intrusion detection," IEEE Commun. Surveys Tuts., vol. 12, no. 3, pp. 343–356, 3rd Quart. 2010.