How to deposit?
A guide to depositing
research data

Creating an account and logging in
Anyone with an active Central Authentication System of the University of Warsaw (CAS) account can deposit research data in the Dane Badawcze UW repository. Using the repository is free for everyone.
To create an account, go to Sign in/Sign up option on the main page of https://danebadawcze.uw.edu.pl/. Then select the Uniwersytet Warszawski as Your institution and confirm with the Log in button.
Please note: the Sign up option should be used only by users who do not have an active CAS UW account. That includes for example researchers from outside the University of Warsaw involved in a project run by the UW or researchers whose affiliation with the UW has ended (they have changed their place of work or completed their doctoral studies).
After confirming with the Log in button, the CAS log-in page will be displayed. Enter ID and password, and use the green button to confirm.
On the next page, fill in the required information, then read and accept the terms of use. Hover your cursor over the tooltip next to any field to learn more about it.
Use the Create account button to save your information. A link will be sent to the e-mail address you have provided, allowing you to verify your account and thus complete the registration process.
Each subsequent log-in will require selecting the Sign in/Sign up button on the main page, then the Log in button below the institution’s name, and entering an ID and password in the CAS system.
Depositing a dataset
Select the Add dataset button on the home page to start depositing a dataset.
Next, select the dataverse of your faculty, and confirm with the Add a dataset to the selected dataverse button. You will be redirected to the page with a metadata form.
Adding metadata
Insert the information about your data into the relevant fields of the metadata form.
Mandatory fields are marked with an asterisk, but filling in as many optional fields as possible is advisable. This will make the data easier to find and interpret.
There are tooltips available next to each metadata field.
It will still be possible to edit and add metadata later on.
Uploading files and selecting a licence
Appropriately prepared data files can be selected from your device and uploaded using the Select files button or dragged and dropped into the indicated widget.
The size limit of a single file is 5 GB, and the size limit for the sum of files being transferred at the same time is 20 GB.
The size of the entire dataset is technically not limited, but it is recommended to not exceed 2 TB.
Recommended file formats: open in a new tab
Each uploaded file must be labelled with an appropriate licence selected by the user from the drop-down menu (the default is a CC0 licence).
Editing the name of each file and adding a short description is also possible.
The repository allows changing licences for multiple files at once. To do so, select the files you want to adjust, then click the Edit button and choose Licence from the drop-down menu.
A new window will pop up – select a desired licence and confirm.
Saving a draft dataset and assigning a DOI identifier
Once the files have been added and the metadata form completed, the dataset needs to be saved using the button at the bottom of the page.
The dataset will be saved as a draft, which means that metadata and files can still be edited at this stage. Therefore, saving a dataset does not mean it becomes public — the draft version is not visible to other users and must be submitted for review before publication. The draft version can also be deleted.
At this stage, the DOI identifier is reserved for the dataset and visible to the depositor. It is not yet active — it will be activated once the administrator publishes the dataset.
Editing a draft dataset
Editing can be enabled on the dataset draft version page. By using the Edit drop-down menu, it is possible to:
- – add more files to the dataset → Files (upload),
- – edit metadata → Metadata,
- – embargo files → Embargo; a functionality allowing to limit the visibility of uploaded files for a certain period after the publication of the dataset – maximum 36 months; embargo can be lifted earlier, but it cannot be restored or extended,
- – create a private URL for an unpublished dataset → Private URL; it can be sent to others, allowing them access to the unpublished dataset; the URL can be deactivated at any time.
The saved draft version can be permanently deleted by selecting the option Delete dataset. This operation is definitive – a deleted draft cannot be restored. Once a dataset has been published, this option will not be available. A published dataset cannot be deleted.
Submitting a dataset for verification
Once ready, submit your dataset for verification using the Submit for review button.
It will be examined by the administrators of the dataverse, who will assess the correctness of the metadata and data, and then either publish the dataset or return it for amendments. A notification will be sent to the email address provided during registration, together with a comment from the administrator.
If corrections are necessary, the dataset has to be edited and resubmitted for review.
If no corrections are needed, the dataset will be published.
Modifying a published dataset
A published dataset cannot be deleted. However, it is possible to modify an existing dataset by creating a new version. Changes can be applied, and the dataset should be resubmitted for verification. The modified dataset will be published as a new version of the existing resource. All previous versions remain available to users – they can be viewed under the Versions tab on the dataset page.
The DOI identifier remains the same for all versions of the dataset. It always resolves to its most recent version.
Guidelines for depositors
The guidelines presented on this page are intended to standardize the rules for depositing data in the repository and to streamline the entire process. Their implementation will facilitate the preparation and submission of data, as well as speed up its verification process.
The introduced standards ensure consistency of the collected resources and accelerate the process of their review and publication. However, taking into account the specific nature and wide diversity of data, in cases where the provisions of these guidelines cannot be applied to a particular dataset or raise doubts, please contact the reviewers directly in order to develop an appropriate solution on an individual basis.
Research data are the materials needed to produce and evaluate the results of scientific research, for example:
– measurement results
– observation notes and literature search results
– annotations to analyzed texts, summaries, and excerpts
– databases of literary motifs, charts, maps, and digital analyses
– questionnaires and survey results
– audio and video recordings, graphics
– mathematical models
– software
– methodological descriptions…
Publications, such as research articles, should be deposited in the UW Institutional Repository.
Research data are made available in the repository together with metadata. These two elements constitute the essential components of a research dataset. Metadata are structured information describing a dataset. They enable research data to be discovered and facilitate their reuse.
Dataset title
– The title should briefly and clearly identify the content of the dataset, explicitly indicate the subject matter of the deposited data, and not contain redundant information, such as author’s name or funding sources.
– The use of phrases such as research data for… or dataset concerning… in the dataset title is not recommended.
– The title of a published research article is not necessarily the best title for a dataset, as the data may be used to prepare further publications.
– Only the first word of the title (and any proper names) should be capitalized.
– The language of the title should be consistent with the language of the remaining metadata. If metadata are created in more than one language, this should be done consistently.
Author
– The metadata fields concerning the author are automatically completed using the information provided in the profile of the user.
– If the author is an individual, the format of the information in the author surname and first name field should be as follows: surname [comma] first name.
– The most commonly used identification system is ORCID. The correct ORCID identifier format is XXXX-XXXX-XXXX-XXXX.
– Information in the affiliation field should be provided in the same language as the remaining metadata.
– The ROR field of the affiliated institution is a suggestion field — after entering part of the institution name, the appropriate name should be selected from the list. If an affiliation is provided, the ROR identifier should also be provided if the institution has such an identifier.
Dataset description
– This field should primarily include information about the deposited files or sets of data files: their content, type, structure, and method of internal organization (especially if ZIP archives are included). Information about the software used to generate the data should be provided, as well as the context of the project and any information that may be useful to potential users, e.g. explanations of abbreviations used.
– It is good practice to create and attach a readme file containing a detailed description of the data. Its preparation may be facilitated by a readme file template. The file should be saved in the simplest possible format, e.g. TXT. The dataset description may indicate the presence of a readme file among the deposited materials, e.g. by adding: Additional explanations are provided in the readme.txt file / For detailed, please consult readme.txt file.
– Information about related publications or grants may be included in the description field only as supplementary information. It should primarily be placed in the appropriate metadata fields (related publication, grant information).
Keywords
– Describing datasets using keywords facilitates the discovery of research outputs.
– Keywords should be linguistically consistent with the rest of the metadata.
– Controlled vocabulary appropriate to the research field may be applied to individual keywords, e.g. medical terminology thesauri or NATO terminology. Dictionaries and databases can be searched, for example, here.
– Each keyword should be entered in a separate field. Keywords entered into a single field should be separated (you may use separators that automatically divide elements separated by commas or semicolons).
Related publication / related dataset
– Only publications / datasets that have already been released / published should be entered. All subfields should be completed.
– The citation field should contain the following information: for publications – author(s), year, title, publisher/journal, journal issue/volume, page numbers, DOI identifier (or another identifier), preferably in the form of a URL. The information should be provided according to the preferred citation standard;
for datasets – author(s), year of publication, title, identifier, repository name, version number.
– The type of relation should be carefully selected and should accurately reflect the relationship between the dataset and the publication.
– If DOI is indicated as the type of publication/dataset identifier, it should be provided in its standard form, not as a URL (e.g. 10.123456/XYZ123 instead of https://doi.org/10.123456/01XYZ123).
– Whenever possible, the URL of the publication/dataset should be based on the persistent DOI identifier: correct – https://doi.org/10.123456/01XYZ123; incorrect – https://strona.wydawcy.pl/tytul-artykulu.
Grant information
– The fields funding institution, institution abbreviation, funding institution ROR, and grant programme are suggestion fields. Suggested information should be used, and unnecessary elements should not be added, e.g. the grant call number: correct – OPUS; incorrect: OPUS 16.
– The appropriate language of metadata in these fields should be ensured — it should be consistent with the remaining metadata.
Language errors
Attention should be paid to correct punctuation, and typos and minor language errors should be eliminated. Examples of frequently occurring errors include: a full stop at the end of the dataset title; missing full stop at the end of the description; inconsistent formatting of keywords; capitalization of words in the dataset title; typographical errors.
File names
They should consist only of the a–Z, 0–9 characters, _ and should not contain spaces, Polish characters, or other special characters. This also applies to files contained within ZIP archives.
The readme file
A documentation element of the dataset; it should be placed at the top of the file list. In most cases, the correct file encoding is UTF-8, and the correct format is TXT. A template may be used and adapted to individual needs.
Versioning
If minor changes are introduced to the metadata of a published dataset, the reviewer will publish a so-called minor version — e.g. v. 1.1, 1.2. If major changes are introduced, especially changes to files, a major version will be published — v. 2, v. 3. Very significant changes to a dataset should be published as a new dataset rather than as a new version of an existing dataset.
Licences
Attention should be paid to the consistent application of licences, e.g. by comparing the information included in the readme file or dataset description with the information provided for individual files. The recommended licences for data files are CC0 or CC BY.
Recommended file formats and preparation of tabular data
Proper preparation of files is crucial to enable correct analysis of tabular data. Information on recommended formats and the proper preparation of tabular data can be found in the document Recommended file formats.
Last updated: June 2026