sparknlp_jsl#

Functions

get_credentials(spark)

Gets John Snow Labs credentials

library_settings(spark)

Gets the library settings

load_license_validator()

pub_version()

Gets the public version of Spark NLP

start([secret, gpu, apple_silicon, aarch64, ...])

Starts a SparkSession with default parameters for Spark NLP for Healthcare.

version()

Gets the version of Spark NLP

get_credentials(spark)#

Gets John Snow Labs credentials

Parameters:

spark (SparkSession) – SparkSession

Returns:

(secretKey, keyId, token)

Return type:

tuple

library_settings(spark)#

Gets the library settings

Parameters:

spark (SparkSession) – SparkSession

Returns:

Library settings

Return type:

str

pub_version()#

Gets the public version of Spark NLP

Returns:

Public version of Spark NLP

Return type:

str

start(secret: str | None = None, gpu: bool = False, apple_silicon: bool = False, aarch64=False, public: str = '', params: dict | None = None, fhir_deid: bool = False)#

Starts a SparkSession with default parameters for Spark NLP for Healthcare.

The default parameters would result in the equivalent of:

SparkSession.builder \
    .appName("Spark NLP Licensed") \
    .master("local[*]") \
    .config("spark.driver.memory", "{{available memory}}") \
    .config("spark.serializer", "org.apache.spark.serializer.KryoSerializer") \
    .config("spark.kryoserializer.buffer.max", "2000M") \
    .config("spark.driver.maxResultSize", "0") \
    .config("spark.extraListeners", "com.johnsnowlabs.license.LicenseLifeCycleManager") \
    .config("spark.jars", "https://pypi.johnsnowlabs.com/<secret>/<licensed jar>") \
    .config("spark.jars.packages", "<open-source Spark NLP Maven coordinate>") \
    .getOrCreate()

The Scala binary version and the Maven/jar artifacts are selected automatically from the installed PySpark version (see the table below), so the same call works on both Spark 3.x (Scala 2.12) and Spark 4.x (Scala 2.13).

Parameters:
  • secret (str, optional) – License secret key. If None, it is read from the SECRET environment variable.

  • gpu (bool, optional) – Whether to use the GPU build of Spark NLP. Defaults to False.

  • apple_silicon (bool, optional) – Whether to use the Apple Silicon (M1/M2) build. Defaults to False.

  • aarch64 (bool, optional) – Whether to use the Linux aarch64 build. Defaults to False.

  • public (str, optional) – Open-source Spark NLP version to use. Defaults to the bundled public version.

  • params (dict, optional) – Extra SparkSession configuration, e.g. {"spark.executor.memory": "8G"}. Defaults to None.

  • fhir_deid (bool, optional) – Whether to use the FHIR De-identification build of Spark NLP for Healthcare. Defaults to False.

Notes

spark.driver.memory defaults to the available system memory.

The artifacts are resolved from the installed PySpark version:

PySpark version

Scala binary

Open-source coordinate

Licensed jar

3.x

2.12

spark-nlp_2.12

spark-nlp-jsl-<v>.jar

4.0.0

2.13

spark-nlp-spark400_2.13

spark-nlp-jsl-spark400_2.13-<v>.jar

4.x above 4.0.0

2.13

spark-nlp_2.13

spark-nlp-jsl_2.13-<v>.jar

Returns:

A SparkSession configured for Healthcare NLP.

Return type:

SparkSession

Raises:

ValueError – If secret is missing.

version()#

Gets the version of Spark NLP

Returns:

Version of Spark NLP

Return type:

str