Introduction to the Kyuubi Configurations System

Kyuubi provides several ways to configure the system and corresponding engines.

Environments

You can configure the environment variables in $KYUUBI_HOME/conf/kyuubi-env.sh, e.g, JAVA_HOME, then this java runtime will be used both for Kyuubi server instance and the applications it launches. You can also change the variable in the subprocess's env configuration file, e.g.$SPARK_HOME/conf/spark-env.sh to use more specific ENV for SQL engine applications.

#!/usr/bin/env bash
#
# Licensed to the Apache Software Foundation (ASF) under one or more
# contributor license agreements.  See the NOTICE file distributed with
# this work for additional information regarding copyright ownership.
# The ASF licenses this file to You under the Apache License, Version 2.0
# (the "License"); you may not use this file except in compliance with
# the License.  You may obtain a copy of the License at
#
#    http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
#
#
# - JAVA_HOME               Java runtime to use. By default use "java" from PATH.
#
#
# - KYUUBI_CONF_DIR         Directory containing the Kyuubi configurations to use.
#                           (Default: $KYUUBI_HOME/conf)
# - KYUUBI_LOG_DIR          Directory for Kyuubi server-side logs.
#                           (Default: $KYUUBI_HOME/logs)
# - KYUUBI_PID_DIR          Directory stores the Kyuubi instance pid file.
#                           (Default: $KYUUBI_HOME/pid)
# - KYUUBI_MAX_LOG_FILES    Maximum number of Kyuubi server logs can rotate to.
#                           (Default: 5)
# - KYUUBI_JAVA_OPTS        JVM options for the Kyuubi server itself in the form "-Dx=y".
#                           (Default: none).
# - KYUUBI_NICENESS         The scheduling priority for Kyuubi server.
#                           (Default: 0)
# - KYUUBI_WORK_DIR_ROOT    Root directory for launching sql engine applications.
#                           (Default: $KYUUBI_HOME/work)
# - HADOOP_CONF_DIR         Directory containing the Hadoop / YARN configuration to use.
#
# - SPARK_HOME              Spark distribution which you would like to use in Kyuubi.
# - SPARK_CONF_DIR          Optional directory where the Spark configuration lives.
#                           (Default: $SPARK_HOME/conf)
#


## Examples ##

# export JAVA_HOME=/usr/jdk64/jdk1.8.0_152
# export HADOOP_CONF_DIR=/usr/ndp/current/mapreduce_client/conf
# export KYUUBI_JAVA_OPTS="-Xmx10g -XX:+UnlockDiagnosticVMOptions -XX:ParGCCardsPerStrideChunk=4096 -XX:+UseParNewGC -XX:+UseConcMarkSweepGC -XX:+CMSConcurrentMTEnabled -XX:CMSInitiatingOccupancyFraction=70 -XX:+UseCMSInitiatingOccupancyOnly -XX:+CMSClassUnloadingEnabled -XX:+CMSParallelRemarkEnabled -XX:+UseCondCardMark -XX:MaxDirectMemorySize=1024m  -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=./logs -verbose:gc -XX:+PrintGCDetails -XX:+PrintGCDateStamps -XX:+PrintTenuringDistribution -Xloggc:./logs/kyuubi-server-gc-%t.log -XX:+UseGCLogFileRotation -XX:NumberOfGCLogFiles=10 -XX:GCLogFileSize=5M -XX:NewRatio=3 -XX:MetaspaceSize=512m"

Kyuubi Configurations

You can configure the Kyuubi properties in $KYUUBI_HOME/conf/kyuubi-defaults.conf. For example:

#
# Licensed to the Apache Software Foundation (ASF) under one or more
# contributor license agreements.  See the NOTICE file distributed with
# this work for additional information regarding copyright ownership.
# The ASF licenses this file to You under the Apache License, Version 2.0
# (the "License"); you may not use this file except in compliance with
# the License.  You may obtain a copy of the License at
#
#    http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
#

## Kyuubi Configurations

#
# kyuubi.authentication           NONE
# kyuubi.frontend.bind.host       localhost
# kyuubi.frontend.bind.port       10009
#

# Details in https://kyuubi.readthedocs.io/en/latest/deployment/settings.html

Authentication

KeyDefaultMeaningTypeSince
kyuubi.authenticationNONEClient authentication types. NOSASL: raw transport. NONE: no authentication check. KERBEROS: Kerberos/GSSAPI authentication. LDAP: Lightweight Directory Access Protocol authentication.string1.0.0
kyuubi.authentication
.ldap.base.dn
<undefined>LDAP base DN.string1.0.0
kyuubi.authentication
.ldap.domain
<undefined>LDAP domain.string1.0.0
kyuubi.authentication
.ldap.url
<undefined>SPACE character separated LDAP connection URL(s).string1.0.0
kyuubi.authentication
.sasl.qop
authSasl QOP enable higher levels of protection for Kyuubi communication with clients. auth - authentication only (default) auth-int - authentication plus integrity protection auth-conf - authentication plus integrity and confidentiality protection. This is applicable only if Kyuubi is configured to use Kerberos authentication. string1.0.0

Backend

KeyDefaultMeaningTypeSince
kyuubi.backend.engine
.exec.pool.keepalive
.time
PT1MTime(ms) that an idle async thread of the operation execution thread pool will wait for a new task to arrive before terminating in SQL engine applicationsduration1.0.0
kyuubi.backend.engine
.exec.pool.shutdown
.timeout
PT10STimeout(ms) for the operation execution thread pool to terminate in SQL engine applicationsduration1.0.0
kyuubi.backend.engine
.exec.pool.size
100Number of threads in the operation execution thread pool of SQL engine applicationsint1.0.0
kyuubi.backend.engine
.exec.pool.wait.queue
.size
100Size of the wait queue for the operation execution thread pool in SQL engine applicationsint1.0.0
kyuubi.backend.server
.exec.pool.keepalive
.time
PT1MTime(ms) that an idle async thread of the operation execution thread pool will wait for a new task to arrive before terminating in Kyuubi serverduration1.0.0
kyuubi.backend.server
.exec.pool.shutdown
.timeout
PT10STimeout(ms) for the operation execution thread pool to terminate in Kyuubi serverduration1.0.0
kyuubi.backend.server
.exec.pool.size
100Number of threads in the operation execution thread pool of Kyuubi serverint1.0.0
kyuubi.backend.server
.exec.pool.wait.queue
.size
100Size of the wait queue for the operation execution thread pool of Kyuubi serverint1.0.0

Delegation

KeyDefaultMeaningTypeSince
kyuubi.delegation.key
.update.interval
PT24Hunused yetduration1.0.0
kyuubi.delegation
.token.gc.interval
PT1Hunused yetduration1.0.0
kyuubi.delegation
.token.max.lifetime
PT168Hunused yetduration1.0.0
kyuubi.delegation
.token.renew.interval
PT168Hunused yetduration1.0.0

Engine

KeyDefaultMeaningTypeSince
kyuubi.engine
.deregister.exception
.classes
A comma separated list of exception classes. If there is any exception thrown, whose class matches the specified classes, the engine would deregister itself.seq1.2.0
kyuubi.engine
.deregister.exception
.messages
A comma separated list of exception messages. If there is any exception thrown, whose message or stacktrace matches the specified message list, the engine would deregister itself.seq1.2.0
kyuubi.engine
.deregister.exception
.ttl
PT30MTime to live(TTL) for exceptions pattern specified in kyuubi.engine.deregister.exception.classes and kyuubi.engine.deregister.exception.messages to deregister engines. Once the total error count hits the kyuubi.engine.deregister.job.max.failures within the TTL, an engine will deregister itself and wait for self-terminated. Otherwise, we suppose that the engine has recovered from temporary failures.duration1.2.0
kyuubi.engine
.deregister.job.max
.failures
4Number of failures of job before deregistering the engine.int1.2.0
kyuubi.engine
.initialize.sql
SHOW DATABASESSemiColon-separated list of SQL statements to be initialized in the newly created engine before queries.string1.2.0
kyuubi.engine.share
.level
USEREngines will be shared in different levels, available configs are: CONNECTION: engine will not be shared but only used by the current client connection USER: engine will be shared by all sessions created by a unique username, see also kyuubi.engine.share.level.sub.domain SERVER: the App will be shared by Kyuubi serversstring1.2.0
kyuubi.engine.share
.level.sub.domain
<undefined>Allow end-users to create a sub-domain for the share level of an engine. A sub-domain is a case-insensitive string values in ^[a-zA-Z_]{1,10}$ form. For example, for USER share level, an end-user can share a certain engine within a sub-domain, not for all of its clients. End-users are free to create multiple engines in the USER share levelstring1.2.0

Frontend

KeyDefaultMeaningTypeSince
kyuubi.frontend
.backoff.slot.length
PT0.1STime to back off during login to the frontend service.duration1.0.0
kyuubi.frontend.bind
.host
<undefined>Hostname or IP of the machine on which to run the frontend service.string1.0.0
kyuubi.frontend.bind
.port
10009Port of the machine on which to run the frontend service.int1.0.0
kyuubi.frontend.login
.timeout
PT20STimeout for Thrift clients during login to the frontend service.duration1.0.0
kyuubi.frontend.max
.message.size
104857600Maximum message size in bytes a Kyuubi server will accept.int1.0.0
kyuubi.frontend.max
.worker.threads
999Maximum number of threads in the of frontend worker thread pool for the frontend serviceint1.0.0
kyuubi.frontend.min
.worker.threads
9Minimum number of threads in the of frontend worker thread pool for the frontend serviceint1.0.0
kyuubi.frontend
.worker.keepalive.time
PT1MKeep-alive time (in milliseconds) for an idle worker threadduration1.0.0

Ha

KeyDefaultMeaningTypeSince
kyuubi.ha.zookeeper
.acl.enabled
falseSet to true if the zookeeper ensemble is kerberizedboolean1.0.0
kyuubi.ha.zookeeper
.connection.base.retry
.wait
1000Initial amount of time to wait between retries to the zookeeper ensembleint1.0.0
kyuubi.ha.zookeeper
.connection.max
.retries
3Max retry times for connecting to the zookeeper ensembleint1.0.0
kyuubi.ha.zookeeper
.connection.max.retry
.wait
30000Max amount of time to wait between retries for BOUNDED_EXPONENTIAL_BACKOFF policy can reach, or max time until elapsed for UNTIL_ELAPSED policy to connect the zookeeper ensembleint1.0.0
kyuubi.ha.zookeeper
.connection.retry
.policy
EXPONENTIAL_BACKOFFThe retry policy for connecting to the zookeeper ensemble, all candidates are: ONE_TIME N_TIME EXPONENTIAL_BACKOFF BOUNDED_EXPONENTIAL_BACKOFF UNTIL_ELAPSEDstring1.0.0
kyuubi.ha.zookeeper
.connection.timeout
15000The timeout(ms) of creating the connection to the zookeeper ensembleint1.0.0
kyuubi.ha.zookeeper
.namespace
kyuubiThe root directory for the service to deploy its instance uri. Additionally, it will creates a -[username] suffixed root directory for each applicationstring1.0.0
kyuubi.ha.zookeeper
.node.creation.timeout
PT2MTimeout for creating zookeeper nodeduration1.2.0
kyuubi.ha.zookeeper
.quorum
The connection string for the zookeeper ensemblestring1.0.0
kyuubi.ha.zookeeper
.session.timeout
60000The timeout(ms) of a connected session to be idledint1.0.0

Kinit

KeyDefaultMeaningTypeSince
kyuubi.kinit.intervalPT1HHow often will Kyuubi server run kinit -kt [keytab] [principal] to renew the local Kerberos credentials cacheduration1.0.0
kyuubi.kinit.keytab<undefined>Location of Kyuubi server's keytab.string1.0.0
kyuubi.kinit.max
.attempts
10How many times will kinit process retryint1.0.0
kyuubi.kinit
.principal
<undefined>Name of the Kerberos principal.string1.0.0

Metrics

KeyDefaultMeaningTypeSince
kyuubi.metrics
.console.interval
PT5SHow often should report metrics to consoleduration1.2.0
kyuubi.metrics
.enabled
trueSet to true to enable kyuubi metrics systemboolean1.2.0
kyuubi.metrics.json
.interval
PT5SHow often should report metrics to json fileduration1.2.0
kyuubi.metrics.json
.location
metricsWhere the json metrics file locatedstring1.2.0
kyuubi.metrics
.prometheus.path
/metricsURI context path of prometheus metrics HTTP serverstring1.2.0
kyuubi.metrics
.prometheus.port
10019Prometheus metrics HTTP server portint1.2.0
kyuubi.metrics
.reporters
JSONA comma separated list for all metrics reporters CONSOLE - ConsoleReporter which outputs measurements to CONSOLE periodically. JMX - JmxReporter which listens for new metrics and exposes them as MBeans. JSON - JsonReporter which outputs measurements to json file periodically. PROMETHEUS - PrometheusReporter which exposes metrics in prometheus format. SLF4J - Slf4jReporter which outputs measurements to system log periodically.seq1.2.0
kyuubi.metrics.slf4j
.interval
PT5SHow often should report metrics to SLF4J loggerduration1.2.0

Operation

KeyDefaultMeaningTypeSince
kyuubi.operation.idle
.timeout
PT3HOperation will be closed when it's not accessed for this duration of timeduration1.0.0
kyuubi.operation
.interrupt.on.cancel
trueWhen true, all running tasks will be interrupted if one cancels a query. When false, all running tasks will remain until finished.boolean1.2.0
kyuubi.operation
.query.timeout
<undefined>Timeout for query executions at server-side, take affect with client-side timeout(java.sql.Statement.setQueryTimeout) together, a running query will be cancelled automatically if timeout. It's off by default, which means only client-side take fully control whether the query should timeout or not. If set, client-side timeout capped at this point. To cancel the queries right away without waiting task to finish, consider enabling kyuubi.operation.interrupt.on.cancel together.duration1.2.0
kyuubi.operation
.scheduler.pool
<undefined>The scheduler pool of job. Note that, this config should be used after change Spark config spark.scheduler.mode=FAIR.string1.1.1
kyuubi.operation
.status.polling
.timeout
PT5STimeout(ms) for long polling asynchronous running sql query's statusduration1.0.0

Session

KeyDefaultMeaningTypeSince
kyuubi.session.check
.interval
PT5MThe check interval for session timeout.duration1.0.0
kyuubi.session.conf
.ignore.list
A comma separated list of ignored keys. If the client connection contains any of them, the key and the corresponding value will be removed silently during engine bootstrap and connection setup. Note that this rule is for server-side protection defined via administrators to prevent some essential configs from tampering but will not forbid users to set dynamic configurations via SET syntax.seq1.2.0
kyuubi.session.conf
.restrict.list
A comma separated list of restricted keys. If the client connection contains any of them, the connection will be rejected explicitly during engine bootstrap and connection setup. Note that this rule is for server-side protection defined via administrators to prevent some essential configs from tampering but will not forbid users to set dynamic configurations via SET syntax.seq1.2.0
kyuubi.session.engine
.check.interval
PT5MThe check interval for engine timeoutduration1.0.0
kyuubi.session.engine
.idle.timeout
PT30Mengine timeout, the engine will self-terminate when it's not accessed for this durationduration1.0.0
kyuubi.session.engine
.initialize.timeout
PT1MTimeout for starting the background engine, e.g. SparkSQLEngine.duration1.0.0
kyuubi.session.engine
.log.timeout
PT24HIf we use Spark as the engine then the session submit log is the console output of spark-submit. We will retain the session submit log until over the config value.duration1.1.0
kyuubi.session.engine
.login.timeout
PT15SThe timeout(ms) of creating the connection to remote sql query engineduration1.0.0
kyuubi.session.engine
.share.level
USER(deprecated) - Using kyuubi.engine.share.level insteadstring1.0.0
kyuubi.session.engine
.spark.main.resource
<undefined>The package used to create Spark SQL engine remote application. If it is undefined, Kyuubi will use the defaultstring1.0.0
kyuubi.session.engine
.startup.error.max
.size
8192During engine bootstrapping, if error occurs, using this config to limit the length error message(characters).int1.1.0
kyuubi.session.idle
.timeout
PT6Hsession idle timeout, it will be closed when it's not accessed for this durationduration1.2.0
kyuubi.session
.timeout
PT6H(deprecated)session timeout, it will be closed when it's not accessed for this durationduration1.0.0

Zookeeper

KeyDefaultMeaningTypeSince
kyuubi.zookeeper
.embedded.client.port
2181clientPort for the embedded zookeeper server to listen for client connections, a client here could be Kyuubi server, engine and JDBC clientint1.2.0
kyuubi.zookeeper
.embedded.client.port
.address
<undefined>clientPortAddress for the embedded zookeeper server tostring1.2.0
kyuubi.zookeeper
.embedded.data.dir
embedded_zookeeperdataDir for the embedded zookeeper server where stores the in-memory database snapshots and, unless specified otherwise, the transaction log of updates to the database.string1.2.0
kyuubi.zookeeper
.embedded.data.log.dir
embedded_zookeeperdataLogDir for the embedded zookeeper server where writes the transaction log .string1.2.0
kyuubi.zookeeper
.embedded.directory
embedded_zookeeperThe temporary directory for the embedded zookeeper serverstring1.0.0
kyuubi.zookeeper
.embedded.election
.port
0electionPort for the embedded zookeeper serverint1.2.0
kyuubi.zookeeper
.embedded.max.client
.connections
120maxClientCnxns for the embedded zookeeper server to limits the number of concurrent connections of a single client identified by IP addressint1.2.0
kyuubi.zookeeper
.embedded.max.session
.timeout
<undefined>maxSessionTimeout in milliseconds for the embedded zookeeper server will allow the client to negotiate. Defaults to 20 times the tickTimeint1.2.0
kyuubi.zookeeper
.embedded.min.session
.timeout
<undefined>minSessionTimeout in milliseconds for the embedded zookeeper server will allow the client to negotiate. Defaults to 2 times the tickTimeint1.2.0
kyuubi.zookeeper
.embedded.port
2181The port of the embedded zookeeper serverint1.0.0
kyuubi.zookeeper
.embedded.quorum.port
0quorumPort for the embedded zookeeper serverint1.2.0
kyuubi.zookeeper
.embedded.server.id
-1serverId for the embedded zookeeper serverint1.2.0
kyuubi.zookeeper
.embedded.tick.time
3000tickTime in milliseconds for the embedded zookeeper serverint1.2.0

Spark Configurations

Via spark-defaults.conf

Setting them in $SPARK_HOME/conf/spark-defaults.conf supplies with default values for SQL engine application. Available properties can be found at Spark official online documentation for Spark Configurations

Via kyuubi-defaults.conf

Setting them in $KYUUBI_HOME/conf/kyuubi-defaults.conf supplies with default values for SQL engine application too. These properties will override all settings in $SPARK_HOME/conf/spark-defaults.conf

Via JDBC Connection URL

Setting them in the JDBC Connection URL supplies session-specific for each SQL engine. For example: jdbc:hive2://localhost:10009/default;#spark.sql.shuffle.partitions=2;spark.executor.memory=5g

  • Runtime SQL Configuration

  • Static SQL and Spark Core Configuration

    • For Static SQL Configurations and other spark core configs, e.g. spark.executor.memory, they will take affect if there is no existing SQL engine application. Otherwise, they will just be ignored

Via SET Syntax

Please refer to the Spark official online documentation for SET Command

Logging

Kyuubi uses log4j for logging. You can configure it using $KYUUBI_HOME/conf/log4j.properties.

#
# Licensed to the Apache Software Foundation (ASF) under one or more
# contributor license agreements.  See the NOTICE file distributed with
# this work for additional information regarding copyright ownership.
# The ASF licenses this file to You under the Apache License, Version 2.0
# (the "License"); you may not use this file except in compliance with
# the License.  You may obtain a copy of the License at
#
#    http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
#

# Set everything to be logged to the console
log4j.rootCategory=INFO, console
log4j.appender.console=org.apache.log4j.ConsoleAppender
log4j.appender.console.target=System.err
log4j.appender.console.layout=org.apache.log4j.PatternLayout
log4j.appender.console.layout.ConversionPattern=%d{yyyy-MM-dd HH:mm:ss.SSS} %p %c{2}: %m%n

Other Configurations

Hadoop Configurations

Specifying HADOOP_CONF_DIR to the directory contains hadoop configuration files or treating them as Spark properties with a spark.hadoop. prefix. Please refer to the Spark official online documentation for Inheriting Hadoop Cluster Configuration. Also, please refer to the Apache Hadoop's online documentation for an overview on how to configure Hadoop.

Hive Configurations

These configurations are used for SQL engine application to talk to Hive MetaStore and could be configured in a hive-site.xml. Placed it in $SPARK_HOME/conf directory, or treating them as Spark properties with a spark.hadoop. prefix.

User Defaults

In Kyuubi, we can configure user default settings to meet separate needs. These user defaults override system defaults, but will be overridden by those from JDBC Connection URL or Set Command if could be. They will take effect when creating the SQL engine application ONLY. User default settings are in the form of ___{username}___.{config key}. There are three continuous underscores(_) at both sides of the username and a dot(.) that separates the config key and the prefix. For example:

# For system defaults
spark.master=local
spark.sql.adaptive.enabled=true
# For a user named kent
___kent___.spark.master=yarn
___kent___.spark.sql.adaptive.enabled=false
# For a user named bob
___bob___.spark.master=spark://master:7077
___bob___.spark.executor.memory=8g

In the above case, if there are related configurations from JDBC Connection URL, kent will run his SQL engine application on YARN and prefer the Spark AQE to be off, while bob will activate his SQL engine application on a Spark standalone cluster with 8g heap memory for each executor and obey the Spark AQE behavior of Kyuubi system default. On the other hand, for those users who do not have custom configurations will use system defaults.