Overview
Using secondary data for your research project can be an efficient and useful method of addressing your research questions.
What is Secondary Data?
Secondary data are defined as data that have been collected for a different purpose (e.g., census data) or to answer a given research question (e.g., collected for the purposes of another research project), which are then made available to be used by another researcher or to answer a different research question.
Secondary Data and Ethical Review
As noted on our ethics overview page, 鈥淏efore you perform any research activity, you must complete an ethics review form via the鈥�痑nd have it reviewed in line with the鈥�University's protocol鈥�.
Please note that this includes all research involving secondary data. The鈥�type of research鈥痓eing carried out will dictate the level of ethics review.
Staff and doctoral student projects using secondary data about humans that contains personally identifiable information and is not in the public domain will require ethics review by the relevant Committee/Body. Undergraduate and master's projects are reviewed by the relevant departmental ethics review process.
Where secondary data about humans is fully anonymised and/or in the public domain, and the potential to do harm is classified as low, staff and doctoral student projects require only a proportionate review at departmental level. Undergraduate and masters projects of this type are reviewed and signed off by the main supervisor, and auto-approved by the , with no additional review required.
Ethical considerations
When working with this type of data, researchers and reviewers are asked to consider the following throughout their ethics application form:
Where did the data come from and how were they originally collected?
Researchers should know how the data were originally obtained, including who collected the data, what methods were used, and whether ethics review was carried out prior to the original data collection. It should be clear that the data were not obtained through unethical, exploitative, or unlawful practices.
Did the original participants provide consent for their data to be used in future research projects?
Secondary data must be used in a way that aligns with the consent originally provided by participants. Participants will have consented for their data to be used for the original purpose, but may have also consented for the data to be used in future research studies including secondary analyses. The data source should be able to provide assurances on participants鈥� original consent.
If participants did not explicitly consent for their data to be used in future research projects, researchers can potentially still use the data providing:
- The dataset is entirely anonymised, and participants cannot be re-identified.
- Re-consenting participants is not feasible or practical in the research context.
- The proposed reuse is proportionate, justified, and low risk.
Does use of the dataset introduce any new ethical risks not anticipated in the original study?
Secondary use may introduce new ethical risks not anticipated in the original study. This includes group-level harms, not just individual privacy risks. Researchers should reflect on whether the data concern sensitive topics, whether findings could stigmatise, criminalise, or disadvantage individuals or groups, and whether new analyses could change how participants or communities are represented.
Additional considerations:
- Would participants reasonably expect their data to be used in this way?
- Will the proposed research and use, management and storage of the data meet with the data source鈥檚 requirements? Have all the appropriate documents been completed and permissions granted?
- How will the data source be acknowledged and referenced?
- Are there any copyright issues with the data?
- How will the data be managed? By combining several data sources, could there be any risk of re-identifying participants?
- How will the data and/or the analysis be presented? Will this continue to ensure the confidentiality and anonymity of participants?
- Will the data identify individuals as being at risk of a condition or disease where they may have otherwise been unaware?