class ChunkEntityResolverModel extends AnnotatorModel[ChunkEntityResolverModel] with ResolverParams with HasStorageModel with HasEmbeddingsProperties with HasCaseSensitiveProperties with Licensed with HasSimpleAnnotate[ChunkEntityResolverModel]

Contains all the parameters to transform a dataset with two Input Annotations of types TOKEN and WORD_EMBEDDINGS, coming from ChunkTokenizer and ChunkEmbeddings Annotators and return the Normalized Entity for a particular trained ontology / curated dataset (e.g. ICD-10, RxNorm, SNOMED etc).

For available pretrained models please see the Models Hub.

Example

Using pretrained models for SNOMED

First the prior steps of the pipeline are defined. Output of types TOKEN and WORD_EMBEDDINGS are needed.

val data = Seq(("A 63-year-old man presents to the hospital ...")).toDF("text")
val docAssembler = new DocumentAssembler().setInputCol("text").setOutputCol("document")
val sentenceDetector = new SentenceDetector().setInputCols("document").setOutputCol("sentence")
val tokenizer = new Tokenizer().setInputCols("sentence").setOutputCol("token")
val word_embeddings = WordEmbeddingsModel.pretrained("embeddings_clinical", "en", "clinical/models")
  .setInputCols("sentence", "token")
  .setOutputCol("word_embeddings")
val icdo_ner = MedicalNerModel.pretrained("ner_bionlp", "en", "clinical/models")
  .setInputCols("sentence", "token", "word_embeddings")
  .setOutputCol("icdo_ner")
val icdo_chunk = new NerConverter().setInputCols("sentence","token","icdo_ner").setOutputCol("icdo_chunk").setWhiteList("Cancer")
val icdo_chunk_embeddings = new ChunkEmbeddings()
  .setInputCols("icdo_chunk", "word_embeddings")
  .setOutputCol("icdo_chunk_embeddings")
val icdo_chunk_resolver = ChunkEntityResolverModel.pretrained("chunkresolve_icdo_clinical", "en", "clinical/models")
  .setInputCols("token","icdo_chunk_embeddings")
  .setOutputCol("tm_icdo_code")
val clinical_ner = MedicalNerModel.pretrained("ner_clinical", "en", "clinical/models")
.setInputCols("sentence", "token", "word_embeddings")
.setOutputCol("ner")
val ner_converter = new NerConverter()
.setInputCols("sentence", "token", "ner")
.setOutputCol("ner_chunk")
val ner_chunk_tokenizer = new ChunkTokenizer()
  .setInputCols("ner_chunk")
  .setOutputCol("ner_token")
val ner_chunk_embeddings = new ChunkEmbeddings()
  .setInputCols("ner_chunk", "word_embeddings")
  .setOutputCol("ner_chunk_embeddings")

Definition of the SNOMED Resolution

val ner_snomed_resolver = ChunkEntityResolverModel.pretrained("chunkresolve_snomed_findings_clinical","en","clinical/models")
    .setInputCols("ner_token","ner_chunk_embeddings").setOutputCol("snomed_result")
val pipelineFull = new Pipeline().setStages(Array(
    docAssembler,
    sentenceDetector,
    tokenizer,
    word_embeddings,

    clinical_ner,
    ner_converter,
    ner_chunk_embeddings,
    ner_chunk_tokenizer,
    ner_snomed_resolver,

    icdo_ner,
    icdo_chunk,
    icdo_chunk_embeddings,
    icdo_chunk_resolver
))
val pipelineModelFull = pipelineFull.fit(data)
val result = pipelineModelFull.transform(data).cache()

Show results

result.selectExpr("explode(snomed_result)")
  .selectExpr(
    "col.metadata.target_text",
    "col.metadata.resolved_text",
    "col.metadata.confidence",
    "col.metadata.all_k_results",
    "col.metadata.all_k_resolutions")
  .filter($"confidence" > 0.2).show(5)
+--------------------+--------------------+----------+--------------------+--------------------+
|         target_text|       resolved_text|confidence|       all_k_results|   all_k_resolutions|
+--------------------+--------------------+----------+--------------------+--------------------+
|hypercholesterolemia|Hypercholesterolemia|    0.2524|13644009:::267432...|Hypercholesterole...|
|                 CBC|             Neocyte|    0.4980|259680000:::11573...|Neocyte:::Blood g...|
|                CD38|       Hypoviscosity|    0.2560|47872005:::370970...|Hypoviscosity:::E...|
|           platelets| Increased platelets|    0.5267|6631009:::2596800...|Increased platele...|
|                CD38|       Hypoviscosity|    0.2560|47872005:::370970...|Hypoviscosity:::E...|
+--------------------+--------------------+----------+--------------------+--------------------+
See also

ChunkEntityResolverApproach on how to train your own model

SentenceEntityResolverModel for sentence level embeddings

Linear Supertypes
HasSimpleAnnotate[ChunkEntityResolverModel], Licensed, HasEmbeddingsProperties, HasStorageModel, HasExcludableStorage, HasStorageReader, HasCaseSensitiveProperties, HasStorageRef, ResolverParams, AnnotatorModel[ChunkEntityResolverModel], CanBeLazy, RawAnnotator[ChunkEntityResolverModel], HasOutputAnnotationCol, HasInputAnnotationCols, HasOutputAnnotatorType, ParamsAndFeaturesWritable, HasFeatures, DefaultParamsWritable, MLWritable, Model[ChunkEntityResolverModel], Transformer, PipelineStage, Logging, Params, Serializable, Serializable, Identifiable, AnyRef, Any
Ordering
  1. Grouped
  2. Alphabetic
  3. By Inheritance
Inherited
  1. ChunkEntityResolverModel
  2. HasSimpleAnnotate
  3. Licensed
  4. HasEmbeddingsProperties
  5. HasStorageModel
  6. HasExcludableStorage
  7. HasStorageReader
  8. HasCaseSensitiveProperties
  9. HasStorageRef
  10. ResolverParams
  11. AnnotatorModel
  12. CanBeLazy
  13. RawAnnotator
  14. HasOutputAnnotationCol
  15. HasInputAnnotationCols
  16. HasOutputAnnotatorType
  17. ParamsAndFeaturesWritable
  18. HasFeatures
  19. DefaultParamsWritable
  20. MLWritable
  21. Model
  22. Transformer
  23. PipelineStage
  24. Logging
  25. Params
  26. Serializable
  27. Serializable
  28. Identifiable
  29. AnyRef
  30. Any
  1. Hide All
  2. Show All
Visibility
  1. Public
  2. All

Instance Constructors

  1. new ChunkEntityResolverModel()
  2. new ChunkEntityResolverModel(uid: String)

    uid

    a unique identifier for the instantiated AnnotatorModel

Type Members

  1. type AnnotationContent = Seq[Row]
    Attributes
    protected
    Definition Classes
    AnnotatorModel
  2. type AnnotatorType = String
    Definition Classes
    HasOutputAnnotatorType

Value Members

  1. final def !=(arg0: Any): Boolean
    Definition Classes
    AnyRef → Any
  2. final def ##(): Int
    Definition Classes
    AnyRef → Any
  3. final def $[T](param: Param[T]): T
    Attributes
    protected
    Definition Classes
    Params
  4. def $$[T](feature: StructFeature[T]): T
    Attributes
    protected
    Definition Classes
    HasFeatures
  5. def $$[K, V](feature: MapFeature[K, V]): Map[K, V]
    Attributes
    protected
    Definition Classes
    HasFeatures
  6. def $$[T](feature: SetFeature[T]): Set[T]
    Attributes
    protected
    Definition Classes
    HasFeatures
  7. def $$[T](feature: ArrayFeature[T]): Array[T]
    Attributes
    protected
    Definition Classes
    HasFeatures
  8. final def ==(arg0: Any): Boolean
    Definition Classes
    AnyRef → Any
  9. def _transform(dataset: Dataset[_], recursivePipeline: Option[PipelineModel]): DataFrame
    Attributes
    protected
    Definition Classes
    AnnotatorModel
  10. def afterAnnotate(dataset: DataFrame): DataFrame
    Attributes
    protected
    Definition Classes
    AnnotatorModel
  11. val allDistancesMetadata: BooleanParam

    whether or not to return an all distance values in the metadata.

    whether or not to return an all distance values in the metadata. Default: False

    Definition Classes
    ResolverParams
  12. val alternatives: IntParam

    number of results to return in the metadata after sorting by last distance calculated

    number of results to return in the metadata after sorting by last distance calculated

    Definition Classes
    ResolverParams
  13. def annotate(annotations: Seq[Annotation]): Seq[Annotation]

    Resolves the ResolverLabel for the given array of TOKEN and WORD_EMBEDDINGS annotations

    Resolves the ResolverLabel for the given array of TOKEN and WORD_EMBEDDINGS annotations

    annotations

    an array of TOKEN and WORD_EMBEDDINGS Annotation objects coming from ChunkTokenizer and ChunkEmbeddings respectively

    returns

    an array of Annotation objects, with the result of the entity resolution for each chunk and the following metadata all_k_results -> Sorted ResolverLabels in the top alternatives that match the distance threshold all_k_resolutions -> Respective ResolverNormalized strings all_k_distances -> Respective distance values after aggregation all_k_wmd_distances -> Respective WMD distance values all_k_tfidf_distances -> Respective TFIDF Cosinge distance values all_k_jaccard_distances -> Respective Jaccard distance values all_k_sorensen_distances -> Respective SorensenDice distance values all_k_jaro_distances -> Respective JaroWinkler distance values all_k_levenshtein_distances -> Respective Levenshtein distance values all_k_confidences -> Respective normalized probabilities based in inverse distance values target_text -> The actual searched string resolved_text -> The top ResolverNormalized string confidence -> Top probability distance -> Top distance value sentence -> Sentence index chunk -> Chunk Index token -> Token index

    Definition Classes
    ChunkEntityResolverModel → HasSimpleAnnotate
  14. final def asInstanceOf[T0]: T0
    Definition Classes
    Any
  15. val auxLabelCol: Param[String]

    Optional column with one extra label per document.

    Optional column with one extra label per document. This extra label will be outputted later on in an additional column (Default: "aux_label".)

  16. val auxLabelMap: StructFeature[Map[String, String]]

    Map[String,String] where key=label and value=auxLabel from a dataset.

  17. def beforeAnnotate(dataset: Dataset[_]): Dataset[_]

    validates the dataset before applying it further down the pipeline

    validates the dataset before applying it further down the pipeline

    Attributes
    protected
    Definition Classes
    ChunkEntityResolverModel → AnnotatorModel
  18. val caseSensitive: BooleanParam
    Definition Classes
    HasCaseSensitiveProperties
  19. final def checkSchema(schema: StructType, inputAnnotatorType: String): Boolean
    Attributes
    protected
    Definition Classes
    HasInputAnnotationCols
  20. final def clear(param: Param[_]): ChunkEntityResolverModel.this.type
    Definition Classes
    Params
  21. def clone(): AnyRef
    Attributes
    protected[lang]
    Definition Classes
    AnyRef
    Annotations
    @throws( ... ) @native()
  22. val confidenceFunction: Param[String]

    what function to use to calculate confidence: INVERSE or SOFTMAX

    what function to use to calculate confidence: INVERSE or SOFTMAX

    Definition Classes
    ResolverParams
  23. def copy(extra: ParamMap): ChunkEntityResolverModel
    Definition Classes
    RawAnnotator → Model → Transformer → PipelineStage → Params
  24. def copyValues[T <: Params](to: T, extra: ParamMap): T
    Attributes
    protected
    Definition Classes
    Params
  25. def createDatabaseConnection(database: Name): RocksDBConnection
    Definition Classes
    HasStorageRef
  26. def createReader(database: Name, connection: RocksDBConnection): WordEmbeddingsReader

    creates WordEmbeddingsReader, based on the DB name and connection

    creates WordEmbeddingsReader, based on the DB name and connection

    database

    Name of the desired database

    connection

    Connection to the RocksDB

    returns

    The instance of the class WordEmbeddingsReader

    Attributes
    protected
    Definition Classes
    ChunkEntityResolverModel → HasStorageReader
  27. val databases: Array[Name]

    This cannot hold EMBEDDINGS since otherwise ER will try to re-save and read embeddings again

    This cannot hold EMBEDDINGS since otherwise ER will try to re-save and read embeddings again

    Attributes
    protected
    Definition Classes
    ChunkEntityResolverModel → HasStorageModel
  28. final def defaultCopy[T <: Params](extra: ParamMap): T
    Attributes
    protected
    Definition Classes
    Params
  29. def deserializeStorage(path: String, spark: SparkSession): Unit
    Definition Classes
    HasStorageModel
  30. def dfAnnotate: UserDefinedFunction
    Definition Classes
    HasSimpleAnnotate
  31. val dimension: IntParam
    Definition Classes
    HasEmbeddingsProperties
  32. val distanceFunction: Param[String]

    what distance function to use for KNN: 'EUCLIDEAN' or 'COSINE'

    what distance function to use for KNN: 'EUCLIDEAN' or 'COSINE'

    Definition Classes
    ResolverParams
  33. val distanceWeights: DoubleArrayParam

    distance weights to apply before pooling: [WMD, TFIDF, Jaccard, SorensenDice, JaroWinkler, Levenshtein]

    distance weights to apply before pooling: [WMD, TFIDF, Jaccard, SorensenDice, JaroWinkler, Levenshtein]

    Definition Classes
    ResolverParams
  34. val enableJaccard: BooleanParam

    whether or not to use Jaccard token distance.

    whether or not to use Jaccard token distance. Default: True

    Definition Classes
    ResolverParams
  35. val enableJaroWinkler: BooleanParam

    whether or not to use Jaro-Winkler character distance.

    whether or not to use Jaro-Winkler character distance. Default: False

    Definition Classes
    ResolverParams
  36. val enableLevenshtein: BooleanParam

    whether or not to use Levenshtein character distance.

    whether or not to use Levenshtein character distance. Default: False

    Definition Classes
    ResolverParams
  37. val enableSorensenDice: BooleanParam

    whether or not to use Sorensen-Dice token distance.

    whether or not to use Sorensen-Dice token distance. Default: False

    Definition Classes
    ResolverParams
  38. val enableTfidf: BooleanParam

    whether or not to use TFIDF token distance.

    whether or not to use TFIDF token distance. Default: True

    Definition Classes
    ResolverParams
  39. val enableWmd: BooleanParam

    whether or not to use WMD token distance.

    whether or not to use WMD token distance. Default: True

    Definition Classes
    ResolverParams
  40. final def eq(arg0: AnyRef): Boolean
    Definition Classes
    AnyRef
  41. def equals(arg0: Any): Boolean
    Definition Classes
    AnyRef → Any
  42. def explainParam(param: Param[_]): String
    Definition Classes
    Params
  43. def explainParams(): String
    Definition Classes
    Params
  44. def extraValidate(structType: StructType): Boolean
    Attributes
    protected
    Definition Classes
    RawAnnotator
  45. def extraValidateMsg: String
    Attributes
    protected
    Definition Classes
    RawAnnotator
  46. final def extractParamMap(): ParamMap
    Definition Classes
    Params
  47. final def extractParamMap(extra: ParamMap): ParamMap
    Definition Classes
    Params
  48. val extramassPenalty: DoubleParam

    penalty for extra words in the knowledge base match during WMD calculation

    penalty for extra words in the knowledge base match during WMD calculation

    Definition Classes
    ResolverParams
  49. val features: ArrayBuffer[Feature[_, _, _]]
    Definition Classes
    HasFeatures
  50. def finalize(): Unit
    Attributes
    protected[lang]
    Definition Classes
    AnyRef
    Annotations
    @throws( classOf[java.lang.Throwable] )
  51. def get[T](feature: StructFeature[T]): Option[T]
    Attributes
    protected
    Definition Classes
    HasFeatures
  52. def get[K, V](feature: MapFeature[K, V]): Option[Map[K, V]]
    Attributes
    protected
    Definition Classes
    HasFeatures
  53. def get[T](feature: SetFeature[T]): Option[Set[T]]
    Attributes
    protected
    Definition Classes
    HasFeatures
  54. def get[T](feature: ArrayFeature[T]): Option[Array[T]]
    Attributes
    protected
    Definition Classes
    HasFeatures
  55. final def get[T](param: Param[T]): Option[T]
    Definition Classes
    Params
  56. def getAllDistancesMetadata: Boolean
    Definition Classes
    ResolverParams
  57. def getAlternatives: Int
    Definition Classes
    ResolverParams
  58. def getAuxLabelCol(): String

    Optional column with one extra label per document.

    Optional column with one extra label per document. This extra label will be outputted later on in an additional column

  59. def getAuxLabelMap(): Map[String, String]

    Map[String,String] where key=label and value=auxLabel from a dataset.

  60. def getCaseSensitive: Boolean
    Definition Classes
    HasCaseSensitiveProperties
  61. final def getClass(): Class[_]
    Definition Classes
    AnyRef → Any
    Annotations
    @native()
  62. def getConfidenceFunction: String
    Definition Classes
    ResolverParams
  63. final def getDefault[T](param: Param[T]): Option[T]
    Definition Classes
    Params
  64. def getDimension: Int
    Definition Classes
    HasEmbeddingsProperties
  65. def getDistanceFunction: String
    Definition Classes
    ResolverParams
  66. def getDistanceWeights: Array[Double]
    Definition Classes
    ResolverParams
  67. def getEnableJaccard: Boolean
    Definition Classes
    ResolverParams
  68. def getEnableJaroWinkler: Boolean
    Definition Classes
    ResolverParams
  69. def getEnableLevenshtein: Boolean
    Definition Classes
    ResolverParams
  70. def getEnableSorensenDice: Boolean
    Definition Classes
    ResolverParams
  71. def getEnableTfidf: Boolean
    Definition Classes
    ResolverParams
  72. def getEnableWmd: Boolean
    Definition Classes
    ResolverParams
  73. def getExtramassPenalty: Double
    Definition Classes
    ResolverParams
  74. def getIncludeStorage: Boolean
    Definition Classes
    HasExcludableStorage
  75. def getInputCols: Array[String]
    Definition Classes
    HasInputAnnotationCols
  76. def getLazyAnnotator: Boolean
    Definition Classes
    CanBeLazy
  77. def getMissAsEmpty: Boolean
    Definition Classes
    ResolverParams
  78. def getNeighbours: Int
    Definition Classes
    ResolverParams
  79. final def getOrDefault[T](param: Param[T]): T
    Definition Classes
    Params
  80. final def getOutputCol: String
    Definition Classes
    HasOutputAnnotationCol
  81. def getParam(paramName: String): Param[Any]
    Definition Classes
    Params
  82. def getPoolingStrategy: String
    Definition Classes
    ResolverParams
  83. def getReader[A](database: Name): StorageReader[A]
    Attributes
    protected
    Definition Classes
    HasStorageReader
  84. def getReturnAllKEmbeddings(): Boolean

    Whether to return all embeddings of all K candidates of the resolution.

    Whether to return all embeddings of all K candidates of the resolution. Embeddings will be in the metadata. Increase in RAM usage to be expected

  85. def getReturnCosineDistances: Boolean

    Whether to calculate and return cosine distances between a chunk/token and the k closest candidates.

    Whether to calculate and return cosine distances between a chunk/token and the k closest candidates. Can improve accuracy but increases computation.

  86. def getSearchTree: SerializableKDTree[TreeData]

    Search Tree.

    Search Tree. Under the hood encapsulates SerializableKDTree. Used to perform the search

  87. def getStorageRef: String
    Definition Classes
    HasStorageRef
  88. def getTermIDF: Map[String, (Int, Double)]

    Inverted Document Frequency of the term.

    Inverted Document Frequency of the term. Used in the TF-IDF method

  89. def getThreshold: Double
    Definition Classes
    ResolverParams
  90. def getUseAuxLabel(): Boolean

    Whether to use Aux Label or not

  91. final def hasDefault[T](param: Param[T]): Boolean
    Definition Classes
    Params
  92. def hasParam(paramName: String): Boolean
    Definition Classes
    Params
  93. def hasParent: Boolean
    Definition Classes
    Model
  94. def hashCode(): Int
    Definition Classes
    AnyRef → Any
    Annotations
    @native()
  95. val includeStorage: BooleanParam
    Definition Classes
    HasExcludableStorage
  96. def initializeLogIfNecessary(isInterpreter: Boolean, silent: Boolean): Boolean
    Attributes
    protected
    Definition Classes
    Logging
  97. def initializeLogIfNecessary(isInterpreter: Boolean): Unit
    Attributes
    protected
    Definition Classes
    Logging
  98. val inputAnnotatorTypes: Array[String]

    Input annotator types: TOKEN, WORD_EMBEDDINGS

    Input annotator types: TOKEN, WORD_EMBEDDINGS

    Definition Classes
    ChunkEntityResolverModel → HasInputAnnotationCols
  99. final val inputCols: StringArrayParam
    Attributes
    protected
    Definition Classes
    HasInputAnnotationCols
  100. final def isDefined(param: Param[_]): Boolean
    Definition Classes
    Params
  101. final def isInstanceOf[T0]: Boolean
    Definition Classes
    Any
  102. final def isSet(param: Param[_]): Boolean
    Definition Classes
    Params
  103. def isTraceEnabled(): Boolean
    Attributes
    protected
    Definition Classes
    Logging
  104. val lazyAnnotator: BooleanParam
    Definition Classes
    CanBeLazy
  105. def log: Logger
    Attributes
    protected
    Definition Classes
    Logging
  106. def logDebug(msg: ⇒ String, throwable: Throwable): Unit
    Attributes
    protected
    Definition Classes
    Logging
  107. def logDebug(msg: ⇒ String): Unit
    Attributes
    protected
    Definition Classes
    Logging
  108. def logError(msg: ⇒ String, throwable: Throwable): Unit
    Attributes
    protected
    Definition Classes
    Logging
  109. def logError(msg: ⇒ String): Unit
    Attributes
    protected
    Definition Classes
    Logging
  110. def logInfo(msg: ⇒ String, throwable: Throwable): Unit
    Attributes
    protected
    Definition Classes
    Logging
  111. def logInfo(msg: ⇒ String): Unit
    Attributes
    protected
    Definition Classes
    Logging
  112. def logName: String
    Attributes
    protected
    Definition Classes
    Logging
  113. def logTrace(msg: ⇒ String, throwable: Throwable): Unit
    Attributes
    protected
    Definition Classes
    Logging
  114. def logTrace(msg: ⇒ String): Unit
    Attributes
    protected
    Definition Classes
    Logging
  115. def logWarning(msg: ⇒ String, throwable: Throwable): Unit
    Attributes
    protected
    Definition Classes
    Logging
  116. def logWarning(msg: ⇒ String): Unit
    Attributes
    protected
    Definition Classes
    Logging
  117. val missAsEmpty: BooleanParam

    whether or not to return an empty annotation on unmatched chunks

    whether or not to return an empty annotation on unmatched chunks

    Definition Classes
    ResolverParams
  118. def msgHelper(schema: StructType): String
    Attributes
    protected
    Definition Classes
    HasInputAnnotationCols
  119. final def ne(arg0: AnyRef): Boolean
    Definition Classes
    AnyRef
  120. val neighbours: IntParam

    number of neighbours to consider in the KNN query to calculate WMD

    number of neighbours to consider in the KNN query to calculate WMD

    Definition Classes
    ResolverParams
  121. final def notify(): Unit
    Definition Classes
    AnyRef
    Annotations
    @native()
  122. final def notifyAll(): Unit
    Definition Classes
    AnyRef
    Annotations
    @native()
  123. def onWrite(path: String, spark: SparkSession): Unit
    Attributes
    protected
    Definition Classes
    HasStorageModel → ParamsAndFeaturesWritable
  124. val outputAnnotatorType: AnnotatorType

    Output annotator types: ENTITY

    Output annotator types: ENTITY

    Definition Classes
    ChunkEntityResolverModel → HasOutputAnnotatorType
  125. final val outputCol: Param[String]
    Attributes
    protected
    Definition Classes
    HasOutputAnnotationCol
  126. lazy val params: Array[Param[_]]
    Definition Classes
    Params
  127. var parent: Estimator[ChunkEntityResolverModel]
    Definition Classes
    Model
  128. val poolingStrategy: Param[String]

    pooling strategy to aggregate distances: AVERAGE or SUM

    pooling strategy to aggregate distances: AVERAGE or SUM

    Definition Classes
    ResolverParams
  129. val readers: Map[Name, StorageReader[_]]
    Attributes
    protected
    Definition Classes
    HasStorageReader
    Annotations
    @transient()
  130. val returnAllKEmbeddings: BooleanParam

    Whether to return all embeddings of all K candidates of the resolution.

    Whether to return all embeddings of all K candidates of the resolution. Embeddings will be in the metadata. Increase in RAM usage to be expected (Default: false)

  131. val returnCosineDistances: BooleanParam

    Whether to calculate and return cosine distances between a chunk/token and the k closest candidates.

    Whether to calculate and return cosine distances between a chunk/token and the k closest candidates. Can improve accuracy but increases computation (Default: true)

  132. def save(path: String): Unit
    Definition Classes
    MLWritable
    Annotations
    @Since( "1.6.0" ) @throws( ... )
  133. def saveStorage(path: String, spark: SparkSession, withinStorage: Boolean): Unit
    Definition Classes
    HasStorageModel
  134. val searchTree: StructFeature[SerializableKDTree[TreeData]]

    Search Tree.

    Search Tree. Under the hood encapsulates SerializableKDTree. Used to perform the search

  135. def serializeStorage(path: String, spark: SparkSession): Unit
    Definition Classes
    HasStorageModel
  136. def set[T](feature: StructFeature[T], value: T): ChunkEntityResolverModel.this.type
    Attributes
    protected
    Definition Classes
    HasFeatures
  137. def set[K, V](feature: MapFeature[K, V], value: Map[K, V]): ChunkEntityResolverModel.this.type
    Attributes
    protected
    Definition Classes
    HasFeatures
  138. def set[T](feature: SetFeature[T], value: Set[T]): ChunkEntityResolverModel.this.type
    Attributes
    protected
    Definition Classes
    HasFeatures
  139. def set[T](feature: ArrayFeature[T], value: Array[T]): ChunkEntityResolverModel.this.type
    Attributes
    protected
    Definition Classes
    HasFeatures
  140. final def set(paramPair: ParamPair[_]): ChunkEntityResolverModel.this.type
    Attributes
    protected
    Definition Classes
    Params
  141. final def set(param: String, value: Any): ChunkEntityResolverModel.this.type
    Attributes
    protected
    Definition Classes
    Params
  142. final def set[T](param: Param[T], value: T): ChunkEntityResolverModel.this.type
    Definition Classes
    Params
  143. def setAllDistancesMetadata(v: Boolean): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  144. def setAlternatives(a: Int): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  145. def setAuxLabelCol(c: String): ChunkEntityResolverModel.this.type

    Optional column with one extra label per document.

    Optional column with one extra label per document. This extra label will be outputted later on in an additional column

  146. def setAuxLabelMap(m: Map[String, String]): ChunkEntityResolverModel.this.type

    Map[String,String] where key=label and value=auxLabel from a dataset.

  147. def setCaseSensitive(value: Boolean): ChunkEntityResolverModel.this.type
    Definition Classes
    HasCaseSensitiveProperties
  148. def setConfidenceFunction(v: String): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  149. def setDefault[T](feature: StructFeature[T], value: () ⇒ T): ChunkEntityResolverModel.this.type
    Attributes
    protected
    Definition Classes
    HasFeatures
  150. def setDefault[K, V](feature: MapFeature[K, V], value: () ⇒ Map[K, V]): ChunkEntityResolverModel.this.type
    Attributes
    protected
    Definition Classes
    HasFeatures
  151. def setDefault[T](feature: SetFeature[T], value: () ⇒ Set[T]): ChunkEntityResolverModel.this.type
    Attributes
    protected
    Definition Classes
    HasFeatures
  152. def setDefault[T](feature: ArrayFeature[T], value: () ⇒ Array[T]): ChunkEntityResolverModel.this.type
    Attributes
    protected
    Definition Classes
    HasFeatures
  153. final def setDefault(paramPairs: ParamPair[_]*): ChunkEntityResolverModel.this.type
    Attributes
    protected
    Definition Classes
    Params
  154. final def setDefault[T](param: Param[T], value: T): ChunkEntityResolverModel.this.type
    Attributes
    protected
    Definition Classes
    Params
  155. def setDimension(value: Int): ChunkEntityResolverModel.this.type
    Definition Classes
    HasEmbeddingsProperties
  156. def setDistanceFunction(value: String): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  157. def setDistanceWeights(v: Array[Double]): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  158. def setEnableJaccard(v: Boolean): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  159. def setEnableJaroWinkler(v: Boolean): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  160. def setEnableLevenshtein(v: Boolean): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  161. def setEnableSorensenDice(v: Boolean): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  162. def setEnableTfidf(v: Boolean): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  163. def setEnableWmd(v: Boolean): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  164. def setExtramassPenalty(emp: Double): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  165. def setIncludeStorage(value: Boolean): ChunkEntityResolverModel.this.type
    Definition Classes
    HasExcludableStorage
  166. final def setInputCols(value: String*): ChunkEntityResolverModel.this.type
    Definition Classes
    HasInputAnnotationCols
  167. final def setInputCols(value: Array[String]): ChunkEntityResolverModel.this.type
    Definition Classes
    HasInputAnnotationCols
  168. def setLazyAnnotator(value: Boolean): ChunkEntityResolverModel.this.type
    Definition Classes
    CanBeLazy
  169. def setMissAsEmpty(v: Boolean): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  170. def setNeighbours(k: Int): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  171. final def setOutputCol(value: String): ChunkEntityResolverModel.this.type
    Definition Classes
    HasOutputAnnotationCol
  172. def setParent(parent: Estimator[ChunkEntityResolverModel]): ChunkEntityResolverModel
    Definition Classes
    Model
  173. def setPoolingStrategy(value: String): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  174. def setReturnAllKEmbeddings(b: Boolean): ChunkEntityResolverModel.this.type

    Whether to return all embeddings of all K candidates of the resolution.

    Whether to return all embeddings of all K candidates of the resolution. Embeddings will be in the metadata. Increase in RAM usage to be expected

  175. def setReturnCosineDistances(value: Boolean): ChunkEntityResolverModel.this.type

    Whether to calculate and return cosine distances between a chunk/token and the k closest candidates.

    Whether to calculate and return cosine distances between a chunk/token and the k closest candidates. Can improve accuracy but increases computation.

  176. def setSearchTree(tree: SerializableKDTree[TreeData]): ChunkEntityResolverModel.this.type

    Search Tree.

    Search Tree. Under the hood encapsulates SerializableKDTree. Used to perform the search

  177. def setStorageRef(value: String): ChunkEntityResolverModel.this.type
    Definition Classes
    HasStorageRef
  178. def setTermIDF(freqs: Map[String, (Int, Double)]): ChunkEntityResolverModel.this.type

    Inverted Document Frequency of the term.

    Inverted Document Frequency of the term. Used in the TF-IDF method

  179. def setThreshold(dist: Double): ChunkEntityResolverModel.this.type
    Definition Classes
    ResolverParams
  180. def setUseAuxLabel(b: Boolean): ChunkEntityResolverModel.this.type

    Whether to use Aux Label or not

  181. val storageRef: Param[String]
    Definition Classes
    HasStorageRef
  182. final def synchronized[T0](arg0: ⇒ T0): T0
    Definition Classes
    AnyRef
  183. val termIDF: StructFeature[Map[String, (Int, Double)]]

    Inverted Document Frequency of the term.

    Inverted Document Frequency of the term. Used in the TF-IDF method

  184. val threshold: DoubleParam

    threshold value for the aggregated distance

    threshold value for the aggregated distance

    Definition Classes
    ResolverParams
  185. def toString(): String
    Definition Classes
    Identifiable → AnyRef → Any
  186. final def transform(dataset: Dataset[_]): DataFrame
    Definition Classes
    AnnotatorModel → Transformer
  187. def transform(dataset: Dataset[_], paramMap: ParamMap): DataFrame
    Definition Classes
    Transformer
    Annotations
    @Since( "2.0.0" )
  188. def transform(dataset: Dataset[_], firstParamPair: ParamPair[_], otherParamPairs: ParamPair[_]*): DataFrame
    Definition Classes
    Transformer
    Annotations
    @Since( "2.0.0" ) @varargs()
  189. final def transformSchema(schema: StructType): StructType
    Definition Classes
    RawAnnotator → PipelineStage
  190. def transformSchema(schema: StructType, logging: Boolean): StructType
    Attributes
    protected
    Definition Classes
    PipelineStage
    Annotations
    @DeveloperApi()
  191. val uid: String
    Definition Classes
    ChunkEntityResolverModel → Identifiable
  192. val useAuxLabel: BooleanParam

    Whether to use Aux Label or not (Default: false)

  193. def validate(schema: StructType): Boolean
    Attributes
    protected
    Definition Classes
    RawAnnotator
  194. def validateStorageRef(dataset: Dataset[_], inputCols: Array[String], annotatorType: String): Unit
    Definition Classes
    HasStorageRef
  195. final def wait(): Unit
    Definition Classes
    AnyRef
    Annotations
    @throws( ... )
  196. final def wait(arg0: Long, arg1: Int): Unit
    Definition Classes
    AnyRef
    Annotations
    @throws( ... )
  197. final def wait(arg0: Long): Unit
    Definition Classes
    AnyRef
    Annotations
    @throws( ... ) @native()
  198. def wrapColumnMetadata(col: Column): Column
    Attributes
    protected
    Definition Classes
    RawAnnotator
  199. def wrapEmbeddingsMetadata(col: Column, embeddingsDim: Int, embeddingsRef: Option[String]): Column
    Attributes
    protected
    Definition Classes
    HasEmbeddingsProperties
  200. def wrapSentenceEmbeddingsMetadata(col: Column, embeddingsDim: Int, embeddingsRef: Option[String]): Column
    Attributes
    protected
    Definition Classes
    HasEmbeddingsProperties
  201. def write: MLWriter
    Definition Classes
    ParamsAndFeaturesWritable → DefaultParamsWritable → MLWritable

Inherited from HasSimpleAnnotate[ChunkEntityResolverModel]

Inherited from Licensed

Inherited from HasEmbeddingsProperties

Inherited from HasStorageModel

Inherited from HasExcludableStorage

Inherited from HasStorageReader

Inherited from HasCaseSensitiveProperties

Inherited from HasStorageRef

Inherited from ResolverParams

Inherited from AnnotatorModel[ChunkEntityResolverModel]

Inherited from CanBeLazy

Inherited from RawAnnotator[ChunkEntityResolverModel]

Inherited from HasOutputAnnotationCol

Inherited from HasInputAnnotationCols

Inherited from HasOutputAnnotatorType

Inherited from ParamsAndFeaturesWritable

Inherited from HasFeatures

Inherited from DefaultParamsWritable

Inherited from MLWritable

Inherited from Model[ChunkEntityResolverModel]

Inherited from Transformer

Inherited from PipelineStage

Inherited from Logging

Inherited from Params

Inherited from Serializable

Inherited from Serializable

Inherited from Identifiable

Inherited from AnyRef

Inherited from Any

Parameters

Annotator types

Required input and expected output annotator types

Members

Parameter setters

Parameter getters