Turn packet captures into flow feature tables, quickly, without falling over on hostile traffic.
Most machine learning on network traffic starts by converting a pcap into per-flow features. That step is usually treated as plumbing, which is a mistake, because it is where two things go wrong. It is slow, and it falls over on exactly the traffic you most want to study.
flowlens reads a capture, groups packets into bidirectional flows, and writes 56 features per flow as CSV. It holds constant memory per flow and bounded memory overall, so a flood that opens millions of one-packet flows does not take it down.
flowlens capture.pcap -o flows.csvOn a synthetic 211 MB capture of 1,000,000 packets: normal TCP conversations with handshakes and teardown, DNS-shaped UDP, and a SYN flood where every packet opens a new 5-tuple that never completes.
| Flow ceiling | Throughput | Peak concurrent flows | Evictions |
|---|---|---|---|
| 1,000,000 (default) | 185,000 pkt/s | 374,154 | 0 |
| 5,000 | 222,000 pkt/s | 5,000 | 370,032 |
Single-threaded, Release build, GCC 16.1, AMD Ryzen. Median of three runs. 405,562 flows extracted, 63 output columns.
The second row is the interesting one. Capping the table made it faster, not slower.
An end-to-end number says how fast the pipeline runs. It does not say what to optimise.
bench/bench_pipeline.cpp isolates each stage over in-memory frames, so file I/O and CSV
writing are excluded:
| Stage | Cost | Rate |
|---|---|---|
| Parse one frame | 12.1 ns | 81.5 M/s |
| Parse + table, 16 concurrent flows | 15.0 M/s | |
| Parse + table, 256 flows | 12.0 M/s | |
| Parse + table, 4,096 flows | 4.6 M/s | |
| Parse + table, 65,536 flows | 2.0 M/s | |
| Parse + table under eviction, 512-flow cap | 1.6 M/s | |
| Welford update | 6.0 ns | 166.8 M/s |
| Build one feature row | 1.24 µs | once per flow |
GCC 16.1 Release, 20 cores at 2.0 GHz, L2 1 MiB per core, L3 16 MiB.
Parsing is not the bottleneck and never was. It runs at 81 M packets/s, roughly 400 times faster than the end-to-end figure. The flow table is the bottleneck, and the reason is cache: throughput falls 7.5× as the working set grows from 16 to 65,536 concurrent flows, with the sharpest drop between 4,096 and 65,536, which is where the table stops fitting in L2 and then L3.
That is the mechanism behind the table above. Capping the flow table does cost eviction bookkeeping, and the benchmark prices it, but under a flood that cost is smaller than what you save by keeping the table cache-resident. Bounding memory turned out to be a throughput win as well as a safety property.
The gap between these numbers and the 185k pkt/s end-to-end figure is file reading and CSV formatting, which the benchmark deliberately excludes. If the goal were raw speed, that is where to look next, not at the parser.
The benchmark above is a flood stress test, not a representative workload: a million packets producing 405,562 flows, averaging 2.5 packets each, with 374,154 open at once. That is the pathological case the flow table is designed to survive.
Real captures look nothing like it. CTU-13 botnet traffic, same binary, same machine, median of three runs each:
| Capture | Packets | Flows | Peak concurrent | Throughput |
|---|---|---|---|---|
| synthetic stress test | 1,000,000 | 405,562 | 374,154 | 136,679 pkt/s |
| CTU-13 sc46 | 4,479,658 | 169,772 | 12,174 | 991,062 pkt/s |
| CTU-13 sc48 | 7,466,160 | 148,621 | 15,973 | 1,540,282 pkt/s |
Real traffic runs 7.3× to 11.3× faster than the synthetic case.
The reason is the same one the flow-ceiling result showed, arriving from the other direction. Real flows are long — 26 and 50 packets each on these captures against 2.5 in the generator — so far fewer are open at once: 12,174 and 15,973 against 374,154, a table 23 to 31 times smaller. It stays in cache, and throughput follows.
So capping the table was not a trick that happened to help a synthetic workload. It was recovering, under adversarial conditions, the locality that real traffic has naturally. That is a better result than the benchmark alone could show, and it is the argument for quoting the 136,679 figure rather than the million-per-second one: the stress test is the number you can rely on, and real traffic is upside.
(The synthetic capture re-measures at 136,679 pkt/s on the current build against the 185,000 recorded above under GCC 16.1. Compare rows within a table, not across them; the three rows here were measured in one session.)
Reproduce with python tools/generate_pcap.py out.pcap 1000000 for the stress
test, and any pcap or pcapng file for the rest.
Keying flows in a hash map is easy. The hard part is knowing when to let go of one.
A flow table that only releases entries on FIN or RST grows without limit under a SYN flood, because none of those flows ever complete. That is not a corner case. It is the traffic you are most likely to be analysing, and it is trivially attacker-controlled.
Three mechanisms bound it:
- Idle timeout. No packet for 120 seconds and the flow is emitted and dropped.
- Hard duration cap. Open for 300 seconds and it is emitted regardless, so a long-lived connection cannot pin an entry forever.
- Capacity ceiling. If the table still exceeds
--max-flows, the least recently seen flows are evicted in batches until it fits.
Expiry runs on a sweep timer driven by capture time, not on every packet. Sweeping per packet would make ingestion cost proportional to table size, which is quadratic over a capture.
When the ceiling does bite, the CLI says so on stderr and the affected flows carry partial features. That is reported rather than hidden, because a silently truncated feature vector is worse than a slow extractor.
Every statistic is folded in as the packet arrives, using Welford's algorithm, so tracking a million-packet flow costs the same as tracking a two-packet one. The obvious alternative, storing inter-arrival times in a vector and reducing at the end, makes memory proportional to packets-per-flow. An attacker chooses how long flows live, so that is the wrong shape.
Welford also holds precision where the textbook sum-of-squares formula does not. There is a test for exactly that, using values offset by 1e9, and another checking that a thousand identical values give a standard deviation of exactly zero rather than NaN.
56 per flow, closely following the CICFlowMeter convention so the output drops into existing pipelines: packet and byte counts per direction, packet length statistics, inter-arrival time statistics for the flow and each direction, TCP flag counts, header lengths, per-second rates, down/up ratio, initial window sizes, and active data packet counts.
Column names come from Flow::feature_names() and the values from Flow::to_row(). A test
asserts the two have the same length, because if they drift apart the CSV silently
mislabels every column, and that is the kind of bug that quietly poisons a downstream
model rather than crashing.
Both capture formats are parsed directly. That costs a few hundred lines and buys a build with no native dependencies, which matters because libpcap is awkward on Windows.
Classic pcap. All four variants: little and big endian, each with microsecond or nanosecond timestamps. A byte-swapped file still parses if you ignore the magic number and just yields nonsense, so the magic is checked explicitly and anything unrecognised is rejected up front.
pcapng. The block-structured successor, and what most modern tools now write by default. Three block types carry what matters: the Section Header fixes byte order, Interface Description Blocks give link type and timestamp resolution, Enhanced Packet Blocks carry packets. Everything else is skipped by its length field.
Timestamp resolution is per interface and is not always microseconds. That is the detail that makes a naive pcapng reader silently wrong: a nanosecond capture read as microseconds stretches every inter-arrival time by a thousand, and every IAT feature with it.
You do not choose. The format is detected from the first four bytes.
Link layers: Ethernet (including stacked VLAN and QinQ tags), raw IP, Linux cooked capture, and BSD loopback.
Captures are routinely taken with a snap length, storing the first N bytes of each frame. That cuts TCP options while leaving the fixed 20-byte header intact, and ports, flags and window all live in those 20 bytes.
flowlens originally refused those packets, reporting bad_header. On a real CTU-13
capture snapped to 54-66 bytes that discarded 48% of the traffic, and the summary
blamed the traffic for what was a capture setting:
| Before | After | |
|---|---|---|
| Flows extracted | 149,183 | 169,772 |
| Undecodable | 2,177,439 (48.6%) | 7,747 (0.17%) |
CTU-13 scenario 46, 4,479,658 packets. The 7,745 remaining are ICMP, correctly unsupported.
The distinction the parser now makes: a length field below its legal minimum is lying and is refused, while one beyond the captured bytes is truncation and is accepted when the fields being read were captured. Those are different failures and deserve different answers.
The CLI reports the breakdown by reason for the same purpose. A capture that is 40% ARP is healthy; one that is 40% too-short means the snap length cut the transport header off and every derived feature would have been wrong. A single total cannot tell you which.
Every offset is bounds-checked before it is read. Packet parsers are a classic source of out-of-bounds reads because the length fields are attacker-controlled: a frame can claim an IHL of 15 while carrying 30 bytes, or an IPv6 extension chain that runs off the end.
There is a test that truncates a valid frame at every single byte offset and asserts none of them parse. The IPv6 extension walk is hop-limited so a chain pointing at itself cannot spin forever. Record lengths above 256 MB are rejected before allocation, because otherwise a corrupt length field asks the allocator for four gigabytes.
The bindings skip CSV entirely. On the 1,000,000-packet capture above:
| Path | Time |
|---|---|
flowlens.extract() |
2.32 s |
| CLI to CSV | 7.03 s |
pandas.read_csv |
2.26 s |
| CSV round trip total | 9.29 s |
4.0x faster, and no 82.6 MB intermediate file. The two paths agree exactly: all 56 feature columns match to the precision the CLI prints.
import flowlens
result = flowlens.extract("capture.pcap")
result.features # (n_flows, 56) float64
result.column("flow_iat_mean")
result.to_torch() # shares the same memory
result.to_pandas()
print(result.stats) # 1,000,000 packets, 405,562 flows, peak 374,154 concurrentThe features have to be computed into memory; that is the work, not a copy.
What is avoided is the second copy, from C++ storage into a fresh NumPy
buffer. Extraction fills one contiguous vector, that vector is moved to the
heap, and NumPy adopts the pointer with a capsule that frees it when the last
reference dies. torch.from_numpy then shares the same block again.
So the matrix is allocated once and never duplicated on the way to a model.
The tests assert this rather than trusting it: features.flags["OWNDATA"] is
False, and a write through the torch tensor is visible in the array.
to_pandas() does copy. Pandas cannot adopt a foreign buffer the way NumPy can.
Build it with -DFLOWLENS_BUILD_PYTHON=ON; see python/README.md.
Needs CMake 3.20 and a C++17 compiler.
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build
./build/flowlens capture.pcap -o flows.csvTests are on by default and fetch GoogleTest automatically:
ctest --test-dir build --output-on-failure # 60 testsSanitizers are opt-in, since MinGW does not ship the runtime libraries:
cmake -S . -B build-asan -DFLOWLENS_SANITIZE=ON -DCMAKE_BUILD_TYPE=DebugBenchmarks fetch Google Benchmark and are off by default:
cmake -S . -B build-bench -DFLOWLENS_BUILD_BENCH=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build-bench && ./build-bench/flowlens_benchCI runs the suite on Linux and Windows across GCC, Clang and MSVC, with a separate sanitizer job on Linux.
The CLI links the GCC runtime statically on MinGW. Windows resolves DLLs from PATH, and
several common installations ship their own older libstdc++-6.dll (Git for Windows is
one). A dynamically linked binary will happily load one of those instead of its own and
then crash inside the standard library with no useful diagnostic. This cost me an
afternoon, so the build no longer depends on what else is installed.
flowlens <capture.pcap> [options]
-o, --output PATH write CSV here (default: stdout)
--idle-timeout N emit a flow after N seconds idle (default: 120)
--max-duration N emit a flow after N seconds open (default: 300)
--max-flows N concurrent flow ceiling, 0 for unlimited (default: 1000000)
-q, --quiet suppress the summary on stderr
Each row carries the 5-tuple, the first-seen timestamp, why the flow ended (fin, reset,
timeout, max_duration, active), and then the features. The end reason is worth
keeping: a capture full of timed-out flows and one full of cleanly-closed flows mean very
different things about the traffic.
- No IP defragmentation. Fragmented packets are parsed if the first fragment carries a complete transport header, and ignored otherwise.
- Not validated against CICFlowMeter yet. The features follow its definitions, but a differential comparison on the same captures has not been run. Until it has, treat the column semantics as "closely following" rather than "identical". This is the next substantial piece of work.
- Single-threaded. Fine at 185k pkt/s for offline analysis. Live capture at line rate would need a different design.
- Two real captures, one malware family. Throughput is now measured on CTU-13 as well as the generator, but both real captures come from the same source. Backbone traffic at a different scale, or a link with many more concurrent flows, would sit somewhere between these numbers and the stress test.
- No live capture. Offline files only.
- Differential validation against CICFlowMeter on shared captures
- Faster CSV writing, now that the benchmarks show it is the remaining cost
MIT. See LICENSE.