While looking into why the sensor in https://docs.stackable.tech/home/stable/airflow/usage-guide/applying-custom-resources/#_dag_code is not able to pull driver logs, I noticed that there is no way to stop driver / executor pods from being immediately deleted after the spark application finishes.
There is a sparkConf value for executors (spark.kubernetes.executor.deleteOnTermination - not considered by the operator and which cannot be acted upon by spark itself because the driver immediately terminates) but nothing comparable for drivers. We could define a new field:
spec:
driver:
ttlSecondsAfterFinished: 600 # default 0 = delete immediately, as today
To enable to sensor use case or allow better debugging. When set to a value > 0, the executor pods should also be kept around as long as deleteOnTermination is not set to false.
While looking into why the sensor in https://docs.stackable.tech/home/stable/airflow/usage-guide/applying-custom-resources/#_dag_code is not able to pull driver logs, I noticed that there is no way to stop driver / executor pods from being immediately deleted after the spark application finishes.
There is a sparkConf value for executors (
spark.kubernetes.executor.deleteOnTermination- not considered by the operator and which cannot be acted upon by spark itself because the driver immediately terminates) but nothing comparable for drivers. We could define a new field:To enable to sensor use case or allow better debugging. When set to a value > 0, the executor pods should also be kept around as long as
deleteOnTerminationis not set to false.