RepeatedKFoldSplitter
Splitter that repeats the K-fold procedure multiple times.
Repeating the folds helps reduce the variance of the performance estimate by averaging over several random resamplings of the same evaluation scheme. This is useful when a single K-fold run is too noisy or when more stable estimates are needed for model comparison.
References
Parameters
- n_splits : integer, default=
5 - Number of folds. Must be an integer between 2 and 20.
- n_repeats : integer, default=
2 - Number of times the K-Fold procedure is repeated. Must be an integer between 2 and 10.
- random_state : integer, default=
42 - Seed used to make the repeated split reproducible.
- test_size : number, default=
0.1 - Proportion of the dataset set aside as a test set. No fold and no hyperparameter search ever sees those rows, so they are scored once by the final model and are the data it can be explained on. Set it to 0 to cross-validate every row, which leaves the run without a test metric and without data to explain, and note that the fold metrics are validation estimates that may carry an optimistic bias if they are used as the final evaluation of the model.
Methods
split_indexes(self, x: 'DashAIDataset', y: 'DashAIDataset') -> 'List[Tuple[List, List]]'
RepeatedKFoldSplitterGenerate train/test index pairs for each repetition of the K-fold split.
Parameters
- x : DashAIDataset
- Input dataset whose length determines the number of available samples.
- y : DashAIDataset
- Target values associated with
x. This argument is accepted for interface consistency but is not used directly by the splitter.
Returns
- list[tuple]
- A list of train/test index pairs for all folds and repeats.
explainable_partitions(cls, split_indexes)
FoldSplitterReturn the partitions of a fold based run an explainer may target.
Parameters
- split_indexes : dict
- The
Run.split_indexespayload, already parsed.
Returns
- dict
- Row indexes for the
trainandtestpartitions.
explainable_splits(cls, split_indexes: 'Dict[str, Any]') -> 'List[Dict[str, Any]]'
BaseSplitterDescribe the partitions of a run that an explainer may target.
Parameters
- split_indexes : dict
- The
Run.split_indexespayload, already parsed.
Returns
- list[dict]
- One
{"name", "rows"}entry per non-empty partition, followed by anallentry. Empty when the run has no data to explain.
get_credential(self, name: str)
ConfigObjectResolve a registered credential component by name.
Parameters
- name : str
- Credential component class name (e.g. "HuggingFaceCredential").
Returns
- BaseCredential
- An instance of the requested credential component.
get_metadata(cls) -> 'dict'
FoldSplitterReturn metadata describing the splitter's compatibility.
get_schema(cls) -> dict
ConfigObjectGenerates the component related Json Schema.
Returns
- dict
- Dictionary representing the Json Schema of the component.
prepare_y(self, y)
BaseSplitterEncode the target variable for stratified splitting.
Parameters
- y : object
- Target values to encode. This may be a list, a pandas-like object, or a DashAI dataset that exposes a single target column.
Returns
- object
- Encoded labels suitable for stratified splitting.
split(self, x: 'DashAIDataset', y: 'DashAIDataset') -> 'Tuple[List[DatasetDict], List[DatasetDict], Dict[str, Any]]'
FoldSplitterCreate folds and return both the partitioned datasets and the indices.
Parameters
- x : DashAIDataset
- Input dataset to split.
- y : DashAIDataset
- Target values associated with
x.
Returns
- tuple[list, list, dict]
- A tuple containing the split datasets for every fold and a mapping from fold names to their corresponding train/test indices.
validate_and_transform(self, raw_data: dict) -> dict
ConfigObjectIt takes the data given by the user to initialize the model and returns it with all the objects that the model needs to work.
Parameters
- raw_data : dict
- A dictionary with the data provided by the user to initialize the model.
Returns
- dict
- A validated dictionary with the necessary objects.