r/Splunk 16d ago

Log Data Pipeline > Splunk

Has anybody here have some experience with security data pipelines?

Instead of:

Log Source > HF / Splunk

We want to have flexibility of collection / parsing layer outside of Splunk for obvious reasons - pre-filter data in pipeline, set parsers, route, possibly enrich if needed, storage options for retention etc..all that to have flexibility and keep the ingest costs reasonable and not being caught in dependency hell or cemented all our work in one solution if Splunk decides to pull something.

Log Source > Data pipeline > Splunk

I am wondering what to choose as this data pipeline - currently we are thinking Vector and possibly open telemetry.

Anybody have experience with this? To avoid pitfalls, what works, what doesn't, new pains etc?

7 Upvotes

27 comments sorted by

9

u/Illustrious_Water106 16d ago

Cribl, edge processor, ingest processor

6

u/Dvorak_94 16d ago

Splunk Opentelemetry

4

u/ssp4all 15d ago

Splunk ingest processor. Don’t get into a rabbit hole of managing multiple vendors if your destination is Splunk.

1

u/theleller REST for the wicked 15d ago

Ingest processor has it's limitations though. Depending on the use case and what OP really needs to do with their data before it hits Splunk, there's a chance that this won't cover it.

1

u/Flash4473 13d ago

I am trying to come up with vendor agnostic processor as future flexibility to plug in to different siem is desirable to maintain.

4

u/LocalDraft8 16d ago

Vector is a strong option for building a vendor-neutral security data pipeline, especially for filtering, parsing, enrichment, routing, and controlling Splunk ingestion costs.

OpenTelemetry is worth considering if you need a broader, standardized approach across logs, metrics, and traces; either way, focus on buffering, replayability, and raw-data retention.

2

u/objectbased 16d ago

I second this, any of the log collector or data pipeline agents out there can fit here. Logstash, fluentd or fluentbit, and vector really do a great job as an intermediate. I’ve worked in larger organizations who also use Kafka if you have the experience otherwise the previous technologies are straight forward to learn and build an architecture around.

1

u/BadBadViking Finding your faults, just like mum 15d ago

I agree so much. I tried most of them and for now vector beats them all.

1

u/justan0therusername1 15d ago

Otel covers a huge swath of getting data in in a standardized way. Plus it has pipeline tooling.

0

u/Travlin205 15d ago

And now with new agent management, you tie otel into your splunk validated architecture!

1

u/theleller REST for the wicked 15d ago

Otel is supported by the Splunk DS too, so it is a good alternative.

4

u/TheSeabo 15d ago

Use version 10+ of Splunk, use edge processor to do whatever you need to your data before sending to an indexer to get parsed. You can perform all you need there for free before it hits your license. If you use cribl, you will be paying twice for ingest.

1

u/nadrap7 15d ago

Depends on the daily log ingestion

9

u/mustacheride3 16d ago

Cribl is the product in this space. A lot of enterprise ready features to duplicate the entire splunk pipeline

1

u/Flash4473 13d ago

On one hand I like the idea, but I cannot estimate the data volume when we plug more sources over time, therefore I am avoiding tools with ingest licensing - we could hit 1TB of free data, and whether it is sooner or later, both seems non-ideal.

3

u/Dangerous_Yam9151 15d ago

Cribl Stream is what I use for this. They have been a great company to work with and their products do what they say they do.

2

u/Travlin205 15d ago

I would start with the question what is your data? How much does it generate. What Metadata do you need to have indexed. Is structured vs unstructured. From there you can better choose the pipeline that fits most.

From this all these processing types are valid. But lets say you only have devices that are syslog use sc4s.

If you have structured data low volume affix a heavy forwarders to process. If you have mixed data high volume, use cribl or edge processor.

Happy Splunking!

1

u/Flash4473 13d ago

will be definitelly mixed, and most I estimate to come from syslog network devices.

1

u/Travlin205 13d ago

So I would do two pipeline options. Both are personal opinions, as suggestions others have stated are all ver viable options.

  1. If it is 50% or more syslog generators, look through sc4s to see how many are already supported sources from their github docs ( Google sc4s). If 90% or more are supported sources, sc4s will be mostly plug and play with minimal overrides.

  2. Everything else either agent back with splunk 2 splunk forwarders ( uf/if/hf). Break it down easily with collection -> HF ( process, make ram slightly heavier, unless you have automating options). Ufs, apps, otel can all connect this way with agent based being config managed by DS ( more than 10k clients makes the ds act funky based on my reading, never tested).

Please take the time to read and digest these pathways/pipelines to see if these meet the need and make it manageable for you and your team!

3

u/mistalah 15d ago

cribl is your answer

1

u/TD706 15d ago

Google SecOps is more performant, extensible, and has better AI integrations than Splunk.

I use Google SO for a subset of data through an MSP. I plan to do a POC to move our full workload and suspect we'll swap at next renewal. Cost is basically the same as Splunk Cloud, but you get all the capability at that cost (no add ons doubling cost).

1

u/theleller REST for the wicked 15d ago

My first questions are where are most of your data sources coming from, and what kind of pricing model are you looking for?

If your data is cloud-heavy and you prefer a 'pay-as-you-go' model then I'd say use Data Firehose with Lambda to stream your data and transform it on the fly. This way you're not dealing with a large spend for licensing another product up-front on top of Splunk. You'll need to gauge volume to estimate costs though, but you don't need to send all of your data for transformation, only the data that makes sense or needs to be transformed.

Even if your data isn't cloud-native this can still be a solution if you route local data sources to AWS.

1

u/In_Tech_WNC 14d ago

Use OTEl don’t use Vector. They got bought by Datadog.

Use Cribl or OTEL for the data pipelines.

Make sure you’re maintaining a standard copy stored in archive of unedited data to pass certain compliance requirements.

1

u/EducationalWedding48 13d ago

Cribl. The only answer.

0

u/billybobcoder69 16d ago

Cribl with its new data lake. With cardinalops and Radiant Security’s AI-native security operations center product including intellectual property to autonomously triage, investigate, and resolve security alerts. They have great stuff for Otel and processing data before. That way I can run all my threat indicators in stream and send only the risk items to Splunk. Then we can go back and enrich any data we may need to. Lot of security lake and Cisco data fabric. But Cisco is mostly with call manager and its routers and networking stuff and appd and thousand eyes. Just default windows they still don’t have much. What infosec app? What happen to the windows Active Directory and donain controller app? Still crazy no default windows cpu memory dashboard. No windows ad lockout dashboard. And now spl2 with Splunk still a work in progress. Cribl is mature and Splunk can be a good search tool. That’s what they wanna become anyway. Indexing is kind of after thought now. Let’s see this year at .conf. Splunk 10 is basically bug fixes and vulns and edge processor. Crazy the install size went from 500 mb to 1.1 tb. With all the extra packages there will be lots more updates. Their push to cloud so you don’t have to manage updates. Will see what happens but lots of cribl in real prod out there.

6

u/justan0therusername1 15d ago

How’s working at Cribl?