The effects of batch size and linger time on Kafka throughput
By default, Kafka attempts to send records as soon as possible, sending up to max.in.flight.requests.per.connection messages per connection. If you attempt to send more than this, the producer will start batching messages, but ultimately if you saturate the connection so there are unacknowledged message batches pending, the producer enters blocking mode.
We can optimise the way the producer batches messages to achieve higher throughput through two settings, batch.size and linger.ms.
linger.ms is the number of milliseconds that the producer waits for more messages before sending a batch. By default, this value is 0, meaning that the producer attempts to send messages immediately. Smaller batches have higher overhead - they are less compressible, and there’s an overhead associated with processing and acknowledging them. Under moderate load, messages may not come frequent enough to fill a batch immediately, but by introducing a small delay (e.g. 50ms), we can increase the likelihood that the producer can batch more messages together, improving throughput.