Optimizing AWS VPC Flow Logs for Network Traffic Analysis with Elasticsearch
For engineers running workloads out of the AWS Sydney region (ap-southeast-2) or the newer Melbourne region (ap-southeast-4), VPC flow logs are often the cheapest source of truth for who is talking to whom across a virtual estate. The records themselves are blunt — accept and reject decisions, byte counts, source and destination pairs — but once they land in Elasticsearch they become a search engine for the network plane. Done poorly, the pipeline floods a cluster, burns through S3 requests, and produces dashboards that time out. Done well, an operator in Perth or Brisbane can pivot across terabytes of telemetry in seconds and spot the misconfigured security group that everyone else missed.
The trick is treating the pipeline as three separate optimisation problems: how the logs are produced and delivered, how they are shaped before indexing, and how the cluster is configured to absorb them. Australian teams working under the Privacy Act 1988, the Notifiable Data Breaches scheme, and for some industries APRA CPS 234 also need to think about retention boundaries and residency, which adds a fourth constraint that offshore templates tend to ignore.
| Pipeline approach | Typical latency | Operational burden | Best fit |
|---|---|---|---|
| Kinesis Firehose to ES | 60–900 s | Low | Steady, predictable flows from many VPCs |
| Lambda to ES bulk API | <30 s | Medium | Small fleets, near-real-time alerting |
| Logstash on EC2/ECS | 5–15 s | High | Heavy parsing, custom enrichment |
| S3 + reindex via EMR | Minutes to hours | Medium | Long-term archival and forensic replay |
Choosing the right flow log format and fields
The default format spits out a space-separated string with version, account, interface ID and a long tail of traffic fields. Switching to a custom format with a fixed field order lets Elasticsearch ingest the data without a fragile regex stage. For most analysis workloads, accept srcaddr, dstaddr, srcport, dstport, protocol, packets, bytes, start, end, action, log-status and vpc-id. Drop the fields you will never query on, like instance-id once it is captured by an enrichment stage, and the document shrinks noticeably.
A common mistake is logging every VPC at the default 10-minute window. For a busy Sydney-based web fleet, that produces millions of records an hour. Tightening the aggregation interval to one minute for sensitive subnets and leaving everything else at ten is usually a better trade — you keep detail where it matters without paying for noise elsewhere.
Delivery and S3 lifecycle tuning
Flow logs destined for Elasticsearch still benefit from landing in S3 first. The bucket acts as a durable buffer, a reindex source, and an evidence store for incident response under the Australian Signals Directorate's Essential Eight guidance. Configure a lifecycle policy that transitions logs to S3 Standard-IA after thirty days, Glacier Instant Retrieval after ninety, and Deep Archive after a year if you are subject to APRA-style record-keeping.
Versioning and server access logging on the bucket should be turned on, and cross-region replication into a second Sydney bucket gives you a clean recovery path without leaving Australian shores. For organisations bound by the Privacy Act, keep the S3 prefix structure aligned to retention windows so that Object Lock or bucket policies can enforce expiry without per-object tagging gymnastics.
Index design and shard sizing
Time-based indices are non-negotiable at scale. Daily indices sized so that the primary shard count sits between 20 and 50 GB is a comfortable default for most flow log volumes; a single busy VPC might warrant hourly indices to keep shard counts manageable. Use ILM to roll hot indices off expensive nodes after a day and merge them into monthly read-only indices that sit on cheaper storage tiers.
Mapping matters more than shard count for query speed. Keyword fields for IP addresses and ports give you term aggregations without the storage overhead of text analysis, while a date type on start and end unlocks time-range queries that return in milliseconds. Disable the _source field only on hot indices where you never need to re-ingest; otherwise keep it on for forensic pivots.
Ingestion with Firehose, Lambda or Logstash
The pipeline comparison above gives a rough steer, but the choice often comes down to who is on call. A small shop in Adelaide with two engineers will probably lean on Firehose because there is nothing to patch at 3 am AEDT. A larger team in Sydney running a managed security service for banks will probably want Logstash on ECS so they can attach custom Grok patterns and call enrichment APIs before indexing. Lambda sits in between: cheap, serverless, but with a hard concurrency ceiling that bites when a noisy neighbour dumps a backlog into the stream.
Whichever path you pick, enable document-level gzip compression on the ES client side and batch to the 5–15 MB range. That single change often halves network egress costs and reduces the number of bulk rejections during traffic spikes.
Query patterns and dashboard performance
Most analysts want three things: a top-talkers dashboard, a rejected-connection panel, and the ability to drill into a specific five-tuple. Build these with composite aggregations on dstaddr and dstport for top-talkers, a terms aggregation on action filtered to REJECT for the security panel, and a saved search template for the five-tuple drill-down. Avoid scripted fields in dashboards; precompute them at ingest time.
For environments with more than fifty VPCs, route each account or VPC into a dedicated index pattern and use cross-cluster search sparingly — it is convenient but adds latency. Cache common aggregations at the dashboard layer and refresh the underlying visualisation on a 30-second interval rather than on every visit; this keeps the cluster cool during business days when most users are at their desks in Melbourne or Brisbane.
Cost control and retention tuning
Elasticsearch is happy to eat your budget if you let it. Force-merge read-only indices down to a single segment to shrink storage, drop replicas on warm tiers, and use data streams with ILM so you never have to remember to delete yesterday's index manually. For audit copies kept outside Elasticsearch, compress flow logs with ZSTD before pushing to S3 — it lands around 70 percent of the gzip size on this kind of repetitive text.
Tag every component — Firehose delivery stream, ES domain, S3 bucket — with the owning business unit and cost centre. AWS Cost Allocation Tags then surface the real per-VPC price of observability, which is usually the number that finally justifies a sharper retention policy to the CFO.
Compliance and operational fit in Australia
Beyond the technical tuning, the operational story matters. Keep all ingestion and storage inside ap-southeast-2 or ap-southeast-4 to satisfy the data sovereignty expectations baked into many Australian government contracts and the IRAP-assessed patterns some agencies require. Document the retention windows against the Privacy Act's APP 11 obligations and your incident response playbook, and make sure the IR team in Canberra or the SOC in Sydney can rehydrate any window of flow logs from S3 within an hour, even if Elasticsearch is rebuilt from scratch.