Hugging Face
Connect datasets from the Hugging Face Hub to Lakehousecat. This enables AI-powered analysis of publicly available or private machine learning datasets.
Connection Fields
| Field | Required | Description |
|---|---|---|
| API Token | Yes | Your Hugging Face personal access token. Generate one at huggingface.co/settings/tokens. |
| Dataset Name | Yes | The dataset identifier in the format owner/dataset-name (e.g., rajpurkar/squad or your-org/private-dataset). |
| Configuration | No | The dataset configuration (subset) to load, if the dataset has multiple configurations. Leave empty to use the default. |
| Split | No | The dataset split to load (e.g., train, test, validation). Leave empty to load all splits. |
Prerequisites
- Admin or Builder role in Lakehousecat
- A Hugging Face account
- For private datasets: a personal access token with read access to the target dataset repository
Notes
- Public datasets on the Hugging Face Hub can be accessed with any valid API token.
- Private or gated datasets require an API token from an account that has been granted access to that specific dataset.
- Hugging Face datasets with multiple splits (e.g.,
train,test) are loaded as separate tables. - Full Load is performed on each data load. For large datasets, configure table and column filters in the Filters tab to limit the scope of semantic extraction.
- Dataset size and load time depend on the dataset size and Hugging Face API response times.
Next Steps
After creating the datasource, go to the Operations tab to trigger semantic extraction.