Describe the bug
The 1.1.6 default of refresh_credentials_interval 18000 (5 hours) is longer than the lifetime of the credentials the plugin itself obtains. When uses assume_role_arn (or assume_role_web_identity_token_file), aws_credentials calls STS without duration_seconds (out_opensearch.rb:238-249), so STS applies its default and issues a 1-hour session. The plugin then waits 5 hours before replacing it. The two defaults are simply inconsistent: for roughly 4 out of every 5 hours the plugin signs every bulk request with a session token it knows the expiry of and has already let lapse.
This is a regression in 1.1.6 specifically. Through 1.1.5 the default was the string "5h" (out_opensearch.rb:198), and because fluentd does not apply :time conversion to default values, the raw string reached timer_execute, where cool.io evaluated "5h".to_f → 5.0 seconds (#130). Credentials were therefore replaced constantly and always fresh. #159 corrected the default to a real 5 hours, which is right in isolation but is the first release where the 5h-refresh/1h-credential mismatch is actually live. Any refresh_credentials_interval above roughly 55 minutes is unsafe given the current 1-hour sessions, and 5 hours is now what users get by default.
out_opensearch_data_stream inherits OpenSearchOutput, so it is affected identically.
To Reproduce
- Create an IAM role that Fluentd can assume, with MaxSessionDuration raised to e.g. 12 hours, and permission to write to an AWS OpenSearch Service domain.
- Install fluent-plugin-opensearch 1.1.6 and run Fluentd to send logs to OpenSearch. Use a role assumption path, not static access_key_id/secret_access_key.
- Confirm records are indexed successfully at startup.
- Keep a steady trickle of records flowing and wait ~65 minutes.
- At the 1-hour mark, all bulk requests begin failing with an expired-token error, and keep failing.
Expected behavior
Default configuration should allow plugin to send the logs seamlessly.
Your Environment
- Fluentd version: 1.19.3
- fluent-plugin-opensearch version: 1.1.6
- Operating system: Alpine Linux 3.23.3
- Kernel version: 6.12.94
Your Configuration
<match **>
@type opensearch_data_stream
<endpoint>
region us-east-1
url https://search-example-xxxxxxxxxxxx.us-east-1.es.amazonaws.com
assume_role_arn arn:aws:iam::123456789012:role/fluentd-opensearch-writer
</endpoint>
logstash_format true
data_stream_name "datastream"
data_stream_template_name "template"
time_key "timestamp"
<buffer>
flush_interval 10s
</buffer>
</match>
Your Error Log
error_class=Fluent::ConfigError error="Failed to create data stream: <datastream> [403] {\"message\":\"The security token included in the request is expired\"}"
Additional context
No response
Describe the bug
The 1.1.6 default of refresh_credentials_interval 18000 (5 hours) is longer than the lifetime of the credentials the plugin itself obtains. When uses assume_role_arn (or assume_role_web_identity_token_file), aws_credentials calls STS without duration_seconds (out_opensearch.rb:238-249), so STS applies its default and issues a 1-hour session. The plugin then waits 5 hours before replacing it. The two defaults are simply inconsistent: for roughly 4 out of every 5 hours the plugin signs every bulk request with a session token it knows the expiry of and has already let lapse.
This is a regression in 1.1.6 specifically. Through 1.1.5 the default was the string "5h" (out_opensearch.rb:198), and because fluentd does not apply :time conversion to default values, the raw string reached timer_execute, where cool.io evaluated "5h".to_f → 5.0 seconds (#130). Credentials were therefore replaced constantly and always fresh. #159 corrected the default to a real 5 hours, which is right in isolation but is the first release where the 5h-refresh/1h-credential mismatch is actually live. Any refresh_credentials_interval above roughly 55 minutes is unsafe given the current 1-hour sessions, and 5 hours is now what users get by default.
out_opensearch_data_stream inherits OpenSearchOutput, so it is affected identically.
To Reproduce
Expected behavior
Default configuration should allow plugin to send the logs seamlessly.
Your Environment
Your Configuration
Your Error Log
error_class=Fluent::ConfigError error="Failed to create data stream: <datastream> [403] {\"message\":\"The security token included in the request is expired\"}"Additional context
No response