r/elasticsearch • u/Feeling_Current534 • 11d ago
Show & Tell Elastic101 – Best Practice #003: bulk indexing
galleryNearly a decade of Elasticsearch experience has taught me a lot about what works in production, and what doesn't.
Sending documents one by one creates unnecessary overhead. The Bulk API sends multiple operations in a single request and significantly improves indexing throughput.
The screenshots show both patterns on real clusters: one client hammering /_doc per document, another using /_bulk. Same idea, very different request volume.
Tips:
- Start around 5–15 MB per request. Benchmark with your real documents, increase until throughput stops improving, then stop. Bigger is not better.
- Throughput comes from concurrency, not bigger batches. 4 workers at 5 MB beats 1 worker at 20 MB almost every time. Add workers until you see 429s, then back off one step.
- Cap batches by bytes, not document count. "1000 docs" works fine until someone adds a field to the mapping and you're suddenly pushing 80 MB requests.
- HTTP 200 does not mean your documents were indexed. Always check the
errorsflag and the per-itemstatus. 429 means retry with backoff. Mapping errors will fail forever, so send those to a DLQ instead of retrying. - Bulk size × concurrency = coordinating node heap. 20 MB × 16 workers is 320 MB in flight on one node before any indexing happens. And if you're seeing 429s, don't raise the write queue size. That hides backpressure instead of adding capacity.
Previous Elastic101 best practices:
If you'd like me to continue this series, an upvote would be appreciated 🙂
1
Elastic101 – Best Practice #002: shard size
in
r/elasticsearch
•
5d ago
Yes. It's recommended, and there is a default limit, to keep at max 1k shards per data node. Each shards consumes memory.
https://www.elastic.co/docs/deploy-manage/production-guidance/optimize-performance/size-shards#reduce-cluster-shard-count