Datasets in a Project#
A project holds datasets under project-local names, in the same way it holds records. Every method here that takes a dataset accepts either the name or the ID.
A dataset in a project is an ordinary dataset. Everything in Datasets applies to it unchanged - entries, specifications, submission, caching, and views all work the same way. What the project adds is the grouping and the name.
Creating datasets#
add_dataset() creates a new dataset in the project. It takes
the same arguments as add_dataset() - see
Adding Datasets - with the dataset type and name required and the rest optional.
>>> ds = proj.add_dataset("singlepoint", "b3lyp barriers")
>>> print(ds.id, ds.name)
377 b3lyp barriers
>>> ds.add_specification("psi4/b3lyp/def2-svp", spec)
>>> ds.add_entry("hooh", hooh_mol)
>>> ds.submit()
The dataset that comes back is a normal dataset of the appropriate type - a
SinglepointDataset here.
Two defaults are inherited from the project unless given explicitly:
default_compute_tag and default_compute_priority. This is the main practical benefit of
creating a dataset inside a project rather than alongside it - the routing settings are set once,
on the project.
>>> proj = client.add_project("Peroxide barrier heights",
... default_compute_tag="big_mem",
... default_compute_priority="low")
>>> ds = proj.add_dataset("optimization", "geometries")
>>> print(ds.default_compute_tag, ds.default_compute_priority)
big_mem PriorityEnum.low
>>> # Or override for this one dataset
>>> ds2 = proj.add_dataset("optimization", "quick geometries",
... default_compute_tag="small_mem",
... default_compute_priority="high")
The name must be unique within the project - reusing a name raises a
PortalRequestError.
Note
Dataset names must be unique both within the project and, for a given dataset type, across the
server. The two checks are separate and existing_ok only affects the second one.
existing_ok=True means “if a dataset of this type and name already exists on the server, add
that dataset to the project instead of creating a new one”. It does not suppress the error
from a name that is already used within this project - that check happens first and is
unconditional. To make a script re-runnable, check
dataset_metadata before adding.
>>> existing = {dm.name for dm in proj.dataset_metadata}
>>> if "b3lyp barriers" not in existing:
... ds = proj.add_dataset("singlepoint", "b3lyp barriers")
... else:
... ds = proj.get_dataset("b3lyp barriers")
Linking existing datasets#
link_dataset() associates a dataset that already exists on
the server with the project. Nothing is copied - the project gains a reference to it.
Only the ID is required. By default the dataset keeps its existing name, description, tagline, and tags.
>>> ds = proj.link_dataset(381)
>>> print(ds.name)
Diatomic geometries
Any of those four can be given to store a project-local value instead.
>>> ds = proj.link_dataset(381,
... name="reference geometries",
... description="Diatomic geometries, used as a reference set",
... tagline="Reference geometries",
... tags=["reference"])
>>> print(ds.name)
reference geometries
Important
These overrides belong to the project, not to the dataset. The dataset’s own name, description, tagline, and tags are unchanged, and the project’s values are substituted in only when the dataset is fetched through the project.
>>> # Through the project - the project's name
>>> print(proj.get_dataset(381).name)
reference geometries
>>> # Directly from the server - the dataset's own name
>>> print(client.get_dataset_by_id(381).name)
Diatomic geometries
A consequence is that the two can drift: renaming the dataset itself with
set_name() does not update the project’s name for it,
and vice versa. The same applies to records, whose project-local name, description, and
tags exist only within the project.
Linking a dataset that is already in this project is an error.
>>> proj.link_dataset(381)
---------------------------------------------------------------------------
PortalRequestError Traceback (most recent call last)
...
PortalRequestError: Request failed: Dataset 381 already linked to project 7 (HTTP status 400)
Getting datasets#
get_dataset() fetches a dataset from the project, by
project-local name or by ID.
>>> ds = proj.get_dataset("b3lyp barriers")
>>> print(ds.id)
377
>>> # By ID works too
>>> ds = proj.get_dataset(377)
>>> print(ds.status())
{'psi4/b3lyp/def2-svp': {<RecordStatusEnum.complete: 'complete'>: 20}}
From here on it is an ordinary dataset - see Using datasets.
Note
Datasets obtained through a project are cached in the same way as those obtained through the
client. If the client was created with a cache_dir, the dataset’s cache file lives there.
See Caching & Views.
Listing datasets#
The dataset_metadata property lists the datasets in the
project without fetching them. Each entry is a
ProjectDatasetMetadata object.
>>> for dm in proj.dataset_metadata:
... print(dm.dataset_id, dm.dataset_type, dm.name, dm.tags)
377 singlepoint b3lyp barriers []
381 optimization reference geometries ['reference']
This is fetched from the server the first time it is accessed and then cached. Call
fetch_dataset_metadata() to refresh it.
>>> proj.fetch_dataset_metadata()
Note that, unlike ProjectRecordMetadata, there is no status
field here - a dataset does not have a single status. Use
status() for a project-wide summary
(Project status), or the dataset’s own
status() for a per-specification breakdown.
Removing datasets#
unlink_datasets() removes datasets from the project. By
default the datasets stay on the server and only the association with the project is dropped.
It accepts a single name or ID, or a list, and the two may be mixed.
>>> proj.unlink_datasets("b3lyp barriers")
>>> proj.unlink_datasets(["reference geometries", 377])
Two flags control how much else goes with it:
delete_datasets- also delete the datasets themselves from the serverdelete_dataset_records- also delete the records held by those datasets
>>> # Remove from the project and delete the dataset, but keep its records
>>> proj.unlink_datasets("b3lyp barriers", delete_datasets=True)
>>> # Remove from the project and delete the dataset and its records
>>> proj.unlink_datasets("b3lyp barriers",
... delete_datasets=True,
... delete_dataset_records=True)
Warning
These deletions are permanent and affect the server, not just the project.
If the dataset is linked into another project as well, delete_datasets=True does not quietly
skip it - it fails with an internal server error, because the dataset is still referenced.
Check with query_project_datasets() first when a dataset may
be shared.
delete_dataset_records=True is best effort in the same way as for records: entries whose
records are also used by another dataset or project are protected by the database and survive,
silently. See Shared records and datasets when deleting.