Short answer: Splunk meters the raw bytes that reach its indexing pipeline, so you reduce Splunk ingest cost by stopping low-value data before that point. Fix noisy audit policies at the source, filter events by ID and pattern at the endpoint, remove fields you don’t search, collapse repeats, and route archive-only telemetry data to cheaper storage. Compression alone does not lower the meter.
Splunk bills climb because every event arrives at full size, including the ones nobody searches. This guide gives you eight steps, in the order we recommend applying them, with configuration examples for Splunk and for NXLog Agent.
What counts toward your Splunk license?
Splunk measures license usage at the indexing pipeline. Splunk’s Admin Manual states that the measured volume is the raw data entering the indexing pipeline, not the compressed data written to disk. However, data filtered and dropped before indexing does not count against your quota. Four more rules on the same page affect your bill:
-
Splunk caps metric events at 150 bytes each for license purposes.
-
Splunk’s own
_internaland_introspectionindexes do not count. -
Summary indexing and metric rollups do not count.
-
Splunk measures usage daily, midnight to midnight, on the license manager’s clock.
Every technique in this guide reduces the data reaching the indexing pipeline. Anything that happens after the meter, such as compression, storage tiering, or retention changes, saves disk and money elsewhere, but not license.
Your pricing model decides how much this matters. Splunk sells activity-based, workload, ingest, and entity pricing. Ingest pricing charges by gigabytes per day, and the steps below hit it directly. Activity-based pricing meters ingest and search activity together, so cutting volume moves one of its two meters. Workload pricing charges for the compute and storage your searches need, so volume reduction still lowers storage and indexer load, but it is not the primary meter. Entity pricing counts monitored hosts, containers, and protected devices, so ingest volume does not move it at all. Splunk does not publish list prices, so we do not quote any. Measure your own volume instead; step 1 shows how.
Where in the pipeline can you cut?
Telemetry data passes through four stages before Splunk indexes it, and you can cut at any of them:
-
The source: The Windows audit policy, the application log level, the firewall’s logging profile.
-
The endpoint agent: The forwarder or collector on the host.
-
An intermediate tier: A Splunk heavy forwarder, Edge Processor, Ingest Processor, or an NXLog Agent relay.
-
The indexer:
props.confandtransforms.confrules, or Ingest Actions rulesets.
The earlier you cut, the more you save, because a dropped event costs no bandwidth, no intermediate-tier CPU, and no indexer time. Each stage limits what you can do, though.
Splunk documents the limits of the endpoint stage.
The universal forwarder does not parse data except in limited situations, and if you need to change or route data based on its contents you need a heavy forwarder.
The universal forwarder can filter Windows Event Log by event code through inputs.conf, but regex routing to the nullQueue runs on a heavy forwarder or indexer, after the data has already crossed the network.
Edge Processor and Ingest Actions, Splunk’s newer options, cover the same use cases at the intermediate tier. In Splunk’s own examples, Edge Processor can match a certain event code, mask the extensive message field at the end of Windows events, and route an unfiltered copy of data to an AWS S3 bucket. The comparison of Splunk’s data management solutions lists the same filter, mask, and Amazon S3 routing capabilities for Ingest Actions. Pipeline templates on Splunk Lantern can filter low-priority Cisco ASA events to reduce license usage. These work, and if you already run them, keep them.
NXLog Agent moves the same work to stage 2. It parses Windows Event Log, Linux and macOS system logs, syslog, and flat files into structured fields on the host, so you can filter on any field, delete fields, deduplicate, and send different subsets to different destinations before anything leaves the machine. The rest of this guide shows how, next to the Splunk-native equivalent where one exists.
Step 1: Measure before you cut
Start with Splunk’s own license accounting.
The license_usage.log file records bytes indexed (b) by source type (st), index (idx), host (h), and source (s).
This search ranks the last 30 days by source type:
index=_internal source=*license_usage.log* type=Usage earliest=-30d@d
| stats sum(b) AS bytes by st
| eval GB=round(bytes/1024/1024/1024, 2)
| sort - GB
| fields st GB
Swap st for h to rank hosts, or for idx to rank indexes.
If host and source values come back empty, you have hit Splunk’s squashing behavior: above a threshold of distinct source/host combinations, Splunk reports only source type and index.
The License Usage Report View in Splunk Web uses the same data.
Then go one level down.
For your top source types, sample events and check len(_raw).
A source type can be expensive because it produces many events, because each event is large, or both.
Each case needs a different fix, which is why the steps below are separate.
If you enrolled your agents in NXLog Platform, the Agents view shows events per second, CPU, and memory per agent, which tells you which hosts to look at first.
Step 2: Turn off noise at the source
The cheapest event is the one Windows never writes. Microsoft’s audit policy documentation rates each subcategory by event volume and recommends a setting per computer type.
The clearest example is Audit Filtering Platform Connection, the subcategory behind events 5156 and 5158. Microsoft rates its volume as high and notes that Success auditing generates one event for every connection made to the system. Its recommendations table has two tiers for each computer type, labeled General and Stronger. The General recommendation for Success auditing is No on domain controllers, member servers, and workstations alike. The Stronger recommendation is IF: enable Success auditing only on high-value hosts where you need to monitor successful inbound or outbound connections to and from untrusted IP addresses. Failure auditing (blocked connections, event 5157) is Yes in both tiers, so keep it on everywhere.
Another documented flood is event 4703. Under Audit Authorization Policy Change, Microsoft warns that applications which adjust token privileges through the AdjustPrivilegesToken API, Configuration Manager being the named example, generate large numbers of 4703 events, and does not recommend Success auditing where such applications run.
Before you filter these downstream, check whether a Group Policy or a third-party agent enabled them in the first place.
The command auditpol /get /category:* on a sample host shows the effective policy.
If nobody needs the events, turn the subcategory off.
If some team does need them, keep the policy and filter at the endpoint instead, which is step 3.
One guardrail from the joint Best Practices for Event Logging and Threat Detection (ASD’s ACSC, CISA, FBI, NSA, and partners, August 2024): keep process creation events with command-line auditing on. Event 4688 can be a large share of Security channel volume on busy hosts, but do not save money there.
Step 3: Filter by event ID at the endpoint
Once the audit policy is right, filter what remains before it leaves the host.
On a Splunk universal forwarder, inputs.conf supports blacklist entries for Windows Event Log inputs, either as plain event codes or as key-value regex pairs.
Splunk’s Windows event log monitoring documentation covers the formats; the inputs.conf configuration reference documents the allow list and deny list formats and the renderXml setting.
[WinEventLog://Security]
disabled = 0
blacklist1 = 5156,5158
blacklist2 = EventCode="4662" Message="Object Type:(?!\s*(domainDNS|groupPolicyContainer|nTDSDSA))"
The second entry keeps 4662 only for the object types your detections use. Step 4 builds the same allowlist on the NXLog side; keep the two in step if you run both.
On NXLog Agent, the Event Log for Windows input module accepts the same structured XML query that Event Viewer uses.
Windows evaluates the query inside the Event Log API, so suppressed events never reach the agent.
Microsoft documents the <Select> and <Suppress> elements and the XPath subset they accept.
The NXLog Agent Event Log for Windows input module reference covers the QueryXML directive.
<Input security_logs>
Module im_msvistalog
<QueryXML>
<QueryList>
<Query Id="0">
<Select Path="Security">*</Select>
<Suppress Path="Security">
*[System[(EventID=5156 or EventID=5158 or EventID=4703)]]
</Suppress>
</Query>
</QueryList>
</QueryXML>
</Input>
A practical tip: build the filter in Event Viewer as a Custom View, switch to the XML tab, and paste the result into QueryXML.
You get a query that a real host has already validated.
On a Splunk heavy forwarder or indexer, the equivalent is a nullQueue transform.
Splunk Cloud customers reach the same outcome through Ingest Actions rulesets or Edge Processor pipelines.
# props.conf
[source::XmlWinEventLog:Security]
TRANSFORMS-drop_wfp = drop_wfp_events
# transforms.conf
[drop_wfp_events]
REGEX = <EventID>(5156|5158)</EventID>
DEST_KEY = queue
FORMAT = nullQueue
The three differ in where the bytes stop.
The forwarder blacklist and the NXLog Agent query stop them on the host.
The nullQueue stops them after they have crossed your network and consumed parsing capacity, but before the license meter.
Step 4: Drop patterns and low-severity events on any source
Event IDs cover Windows. Most other noise has no ID: debug-level application logs, load balancer health checks in web server logs, informational syslog from network gear.
NXLog Agent’s drop() procedure discards the record being processed, and you can call it after any condition on any parsed field.
NXLog’s event filtering guide uses syslog severity as the example.
parse_syslog() sets a normalized $SeverityValue, so one line keeps warnings and above:
<Extension syslog>
Module xm_syslog
</Extension>
<Input udp_listen>
Module im_udp
ListenAddr 0.0.0.0:514
<Exec>
parse_syslog();
# Keep WARNING (3) and above; drop DEBUG and INFO
if $SeverityValue < 3 drop();
</Exec>
</Input>
<Input nginx_access>
Module im_file
File '/var/log/nginx/access.log'
<Exec>
# Load balancer probes carry no security value
if $raw_event =~ /"GET \/healthz / drop();
</Exec>
</Input>
The same pattern lets you filter by field rather than by event ID, which matters for events that are noisy and valuable at the same time. Event 4662 on a domain controller shows why: a few object types drive specific detections, while the rest is volume. Rather than suppress the whole ID, keep what your detection rules reference:
<Input security_logs>
Module im_msvistalog
<QueryXML>
<QueryList>
<Query Id="0">
<Select Path="Security">*</Select>
</Query>
</QueryList>
</QueryXML>
<Exec>
# Keep 4662 only for object types your detections use
if $EventID == 4662 and \
$Message !~ /Object Type:\s+(domainDNS|groupPolicyContainer|nTDSDSA)/ drop();
</Exec>
</Input>
For 4662 the readable object type appears in the rendered message, so this example matches on $Message.
The Event Log for Windows input module also parses the event’s EventData items into fields, so where a value arrives in plain form, such as $SubjectUserName or $ProcessName, compare the field directly.
Adjust the allowlist to match the rules in your Splunk detection content.
Filter on what the detection uses, not on the event ID alone.
Step 5: Trim the event, not only the event count
Because Splunk meters bytes, you can keep every event and still pay less by sending less of each one. Two things inflate Windows events in particular: the rendered description text at the end of each message, and the bookkeeping fields the collector adds.
Splunk documents the first problem.
Its Edge Processor examples name masking "the extensive message field at the end of Windows events" as a use case, Ingest Actions offers the same masking, and the universal forwarder has suppress_text and related settings for the same purpose.
The description is large. NXLog’s event trimming guide publishes a sample 4769 (Kerberos service ticket) message before and after removing the boilerplate. We counted the bytes: the full message is 1,123 bytes, the trimmed version is 603, so the three explanatory paragraphs are 520 bytes, or 46 percent of the event. Sizes vary by event ID and rendering, so run the same count on your own samples, but the pattern holds: the text that explains what a 4769 is repeats identically in every 4769.
NXLog Agent gives you three levels of trimming:
- Cut the boilerplate, keep the message
-
A regex on
$Messageremoves the explanatory paragraphs and leaves the structured summary analysts read. - Delete fields you don’t search
-
The Rewrite extension module’s
Delete,Keep, andRenamedirectives remove fields by name or keep only an allowlist. - Send fields, not text
-
With
to_json()the record’s remaining fields become the payload, so every deleted field is a byte you don’t pay for.
Here is a Security channel collector that applies all three and sends the result to Splunk’s HTTP Event Collector. The HEC output follows NXLog’s Splunk integration guide; the event object is what Splunk indexes, while time, host, source, and sourcetype travel as HEC metadata keys as described in Splunk’s HEC event format documentation.
<Extension json>
Module xm_json
</Extension>
<Extension trim_windows>
Module xm_rewrite
# Collector bookkeeping and rarely searched system fields
Delete EventReceivedTime, SourceModuleName, SourceModuleType, \
Keywords, OpcodeValue, TaskValue, Version, ProviderGuid, \
ExecutionProcessID, ExecutionThreadID
</Extension>
<Extension hec_envelope>
Module xm_rewrite
Keep time, host, source, sourcetype, event
</Extension>
<Input security_logs>
Module im_msvistalog
<QueryXML>
<QueryList>
<Query Id="0">
<Select Path="Security">*</Select>
<Suppress Path="Security">
*[System[(EventID=5156 or EventID=5158 or EventID=4703)]]
</Suppress>
</Query>
</QueryList>
</QueryXML>
<Exec>
# 1. Remove the explanatory paragraphs, keep the structured summary
if $EventID == 4688
{
$Message =~ s/\s*Token Elevation Type indicates the type of.*$//s;
}
else if $EventID == 4769
{
$Message =~ s/\s*This event is generated every time access is.*$//s;
}
# 2. Delete fields you don't search
trim_windows->process();
</Exec>
</Input>
<Output splunk_hec>
Module om_http
URL https://splunk.example.com:8088/services/collector/event
AddHeader Authorization: Splunk 00000000-0000-0000-0000-000000000000
HTTPSCAFile %CERTDIR%/splunk-ca.pem
BatchMode multiline
Compression gzip
<Exec>
# 3. The remaining fields become the event payload
$event = to_json();
# Epoch seconds with microsecond fraction for the HEC "time" key
$time = string(integer($EventTime));
$time =~ /^(?<sec>\d+)(?<ms>\d{6})$/;
$time = $sec + "." + $ms;
$sourcetype = "_json";
$host = $Hostname;
$source = $Channel;
hec_envelope->process();
to_json();
</Exec>
</Output>
<Route security_to_splunk>
Path security_logs => splunk_hec
</Route>
Decide field by field with your detection engineers.
Keywords, OpcodeValue, and Version rarely appear in a search; SubjectUserName, ProcessName, and IpAddress appear constantly.
If a rule or a dashboard references a field, it stays.
Step 6: Collapse repeats
Some sources emit the same event hundreds of times in a row: a service that logs a failed connection every second, a firewall that reports the same denied flow, a misconfigured scheduled task. Each copy costs the same as the first.
NXLog Agent’s De-Duplicator processor module compares each event against the one before it on the fields you choose.
It forwards the first, drops the duplicates, and emits one summary event reading last message repeated n times, the same behavior classic syslog daemons use.
Insert it as a processor in the route:
<Processor dedup>
Module pm_norepeat
# Compare on content, not on timestamps
CheckFields Hostname, SourceName, Message
</Processor>
<Route syslog_to_splunk>
Path udp_listen => dedup => splunk_hec
</Route>
Two things to know before you rely on it.
The module waits one second for duplicates to arrive, so it collapses bursts inside that window rather than an unbounded run, and a source that repeats every few seconds still sends every copy.
And CheckFields matters: left unset it compares $Message alone, which merges identical messages from different hosts, so set it explicitly as above.
NXLog is also phasing this module out. It still ships and still works, but its documentation marks it for removal in a future release and points to module variables as the way to build the same thing: hold the previous event’s key in a module variable, compare the current event against it, and count the repeats yourself. For a configuration you expect to run for years, build it that way. To get a quick measurement today, the De-Duplicator processor module is still the quickest option.
For high-volume numeric telemetry such as performance counters, remember the 150-byte cap on metric events in Splunk’s license accounting. If a source only ever feeds a chart, sending it as metrics rather than events changes what it costs.
Step 7: Route archive-only telemetry data around Splunk
Some telemetry data you must keep but rarely search: compliance retention, full firewall allow logs, verbose application logs you want for incident reconstruction. The joint ASD/CISA guidance recommends exactly this split. It describes a centralized logging facility such as a secured data lake with select, processed telemetry data forwarded to the SIEM, hot and cold storage tiers, and it states that organizations should consider filtering event logs before sending them to a SIEM to limit cost and capacity issues.
NXLog Agent implements the split with two outputs on one input.
Each output receives its own copy of the record, so a drop() in one output does not affect the other.
The Splunk output keeps the detection set; the archive output, here the Amazon S3 output module, keeps everything in JSON.
# Uses the json and hec_envelope extensions and the security_logs input from step 5
<Output splunk_hec>
Module om_http
URL https://splunk.example.com:8088/services/collector/event
AddHeader Authorization: Splunk 00000000-0000-0000-0000-000000000000
HTTPSCAFile %CERTDIR%/splunk-ca.pem
BatchMode multiline
Compression gzip
<Exec>
# Only the events your detection content uses go to Splunk
if not ($EventID in (1102, 4624, 4625, 4648, 4672, 4688, 4698, \
4720, 4728, 4732, 4756, 4768, 4769, 4776)) drop();
$event = to_json();
$time = string(integer($EventTime));
$time =~ /^(?<sec>\d+)(?<ms>\d{6})$/;
$time = $sec + "." + $ms;
$sourcetype = "_json";
$host = $Hostname;
$source = $Channel;
hec_envelope->process();
to_json();
</Exec>
</Output>
<Output archive_s3>
Module om_amazons3
Region us-east-1
Bucket seclog-archive
# Server sets the object path prefix inside the bucket, not a hostname
Server dc01
# Omit the keys to use credentials from the environment, STS, a profile, or instance metadata
AccessKey <YOUR_ACCESS_KEY>
SecretKey <YOUR_SECRET_KEY>
Exec to_json();
</Output>
<Route security_to_splunk_and_s3>
Path security_logs => splunk_hec, archive_s3
</Route>
The event ID list is a starting point, not a recommendation. Derive yours from the rules and dashboards you run, and review it whenever detection content changes.
The archive stays reachable. Splunk Cloud Platform’s Federated Search for Amazon S3 searches S3 data without ingesting it. Splunk positions it for low-frequency, ad hoc searches rather than real-time detection, it requires an AWS Glue Data Catalog, and the validated architecture refers to a per-search scan cost, so check the current terms. NXLog Platform’s log storage is another home for the archive tier, with its own search and retention controls.
Step 8: Batch and compress, but know what it saves
Compression and batching reduce bandwidth and connection overhead. They do not reduce your Splunk license, because Splunk measures raw bytes before compression. Splunk states this in its licensing documentation, and it still trips people up.
Once you forward from many hosts over WAN links or into Splunk Cloud, the bandwidth savings are real. The HTTP(s) output module supports gzip compression and batching, and NXLog’s Splunk integration guide uses both against HEC:
<Output splunk_hec>
Module om_http
URL https://splunk.example.com:8088/services/collector/event
AddHeader Authorization: Splunk 00000000-0000-0000-0000-000000000000
HTTPSCAFile %CERTDIR%/splunk-ca.pem
BatchMode multiline
Compression gzip
</Output>
For agent-to-relay hops inside your own network, the NXLog Transport output module does the same job between NXLog Agent instances; the bandwidth usage guide covers the options.
Where NXLog Platform fits
Everything above is a configuration file. The operational problem is applying it to hundreds or thousands of hosts and confirming that it took effect.
NXLog Platform manages the agent fleet centrally. You build a configuration once, and assign it to agents by group, so a new suppression rule reaches every domain controller in one change, without touching each host. The agent management views show which agents run which configuration, their status, and per-agent throughput, which is how you confirm a filter change produced the drop you expected.
The licensing model also matters here. We license NXLog Platform per data source, not per gigabyte, so trimming and filtering change nothing on the NXLog side of the bill. The only meter your pipeline work moves is Splunk’s.
What not to filter
A filter that removes evidence saves nothing. Three rules limit the risk.
- Filter to your detection content, not to your budget
-
The ASD/CISA guidance defines log quality as the types of events collected, not how well they are formatted, and lists priority sources for enterprise networks: critical systems and data holdings, internet-facing services, identity and domain management servers, edge devices, administrative workstations, and privileged infrastructure such as CI/CD and secrets management. Filters on those sources deserve a second reviewer. The same guidance recommends keeping process creation and command-line auditing on.
- Keep a raw copy somewhere cheap
-
Step 7 exists so that a filter mistake is recoverable. The guidance notes that incidents can take up to 18 months to discover, so set the archive’s retention to cover that window, not the SIEM’s.
- Change one thing, measure, then change the next
-
Apply each filter to a pilot group, rerun the step 1 search after 24 hours, and diff the results. NXLog Platform lets you assign a configuration to a small group first, so the pilot is a selector change rather than a separate deployment. Document why each filter exists; the person who inherits your configuration will need it.
Which step to apply first?
The table compares each technique by where it runs, how much license volume it saves, the risk to detection coverage, and the effort to apply it.
| Technique | Runs on | License impact | Detection risk | Effort |
|---|---|---|---|---|
Fix audit policy (step 2) |
Windows host |
High where WFP or 4703 auditing is on |
Low if you follow Microsoft’s recommendations |
Low |
Event ID filtering (step 3) |
Endpoint agent |
High on Windows Security |
Low for documented noise IDs, higher for broad lists |
Low |
Field-based |
Endpoint agent |
Medium to high on verbose sources |
Medium; filters must map to detection content |
Medium |
Trim message and fields (step 5) |
Endpoint agent |
Medium, applies to every event kept |
Low if searched fields stay |
Medium |
Deduplicate repeats (step 6) |
Endpoint agent or relay |
Source-dependent |
Low; summary event preserves the count |
Low |
Route archive tier elsewhere (step 7) |
Endpoint agent or relay |
High on archive-only sources |
Low with an archive in place |
Medium |
Batch and compress (step 8) |
Any hop |
None |
None |
Low |
|
Heavy forwarder or indexer |
High |
Low |
Low, but bandwidth is already spent |
Start with the meter
Splunk’s license counts bytes at one point in the pipeline. Every step in this guide gets you to that point with fewer of them, in an order that protects detection coverage: policy first, event IDs second, fields and duplicates third, routing fourth. The Splunk-native tools handle part of the job after the data has traveled; NXLog Agent handles all of it on the host, and NXLog Platform pushes the change to the fleet.
To test this against your own volume, start a free NXLog Platform account, point one agent at a noisy domain controller, and compare the step 1 numbers a day later. Our Splunk integration guide has the HEC setup, and we’re happy to help with the field list if you get in touch.
FAQ
- Does filtering at the forwarder reduce Splunk license usage?
-
Yes. Splunk measures volume at the indexing pipeline, and its Admin Manual states that data filtered and dropped before indexing does not count against the license quota. The saving is the same whether the filter runs on a universal forwarder, NXLog Agent, an Edge Processor, or an indexer
nullQueue; the difference is bandwidth and processing spent before the drop. - Does compressing data before sending it to Splunk reduce license usage?
-
No. Splunk measures the raw data entering the indexing pipeline, not the compressed size. Compression reduces network and storage costs only.
- Do Splunk’s internal logs count toward my license?
-
No. Splunk excludes data indexed into
_internaland_introspection, as well as summary indexes and metric rollups. - What is the difference between Ingest Actions, Edge Processor, and Ingest Processor?
-
All three filter, mask, and route data before indexing. Ingest Actions is a UI over
props.confandtransforms.confrulesets on indexers or heavy forwarders. Edge Processor is customer-hosted and uses SPL2 pipelines. Ingest Processor is the Splunk-hosted version of the same pipeline model. All of them work after data has left the endpoint. - Can the Splunk universal forwarder filter events?
-
Partly. It can filter Windows Event Log inputs by event code and by key-value regex in
inputs.conf. Splunk’s documentation states that it does not parse data except in limited situations, so content-based filtering of other sources needs a heavy forwarder or an agent that parses on the host. - Is any of this worthwhile on workload or activity-based pricing?
-
Less directly on workload pricing, which charges for compute and storage rather than gigabytes, so a smaller ingest lowers storage and indexer load without changing the primary meter. Activity-based pricing meters ingest and search together, so cutting volume moves one of the two directly. The detection-value hygiene applies on every model.
- How much will I save?
-
We don’t publish a percentage because the answer depends on your audit policy, your source mix, and how much Windows description text you index today. Run the step 1 search before you apply any filter and again a full day after, then compare the two results.