Overview
The Health Data Science Dissertation involves in-depth research that applies your Graduate Diploma in Health Data Science knowledge to a real-world project. The primary outcome is a practical output, such as a preliminary academic manuscript, an analysis integrated into a larger research initiative, a knitr report designed to facilitate efficient and reproducible reporting, a Shiny web application, or other form of data visualisation. This research is conducted independently, with guidance from academic supervisor(s).
Dissertation courses
Students taking the dissertation can enroll under three different course codes, which distinguish the Units of Credit (UoC) for a given term, 6, 12, or 18. The enrolled UoC determines the time commitment over the course of a twelve-week term, as outlined below.
| HDAT9900 |
6 |
~1.7 days per week |
| HDAT9901 |
12 |
~3.3 days per week |
| HDAT9902 |
18 |
~5 days per week |
Scheduling
The dissertation course comprises 24 UoC in total which be divided in any combination of 6, 12, or 18 UoC as outlined in the table below.
| 2 |
6 UoC |
18 UoC |
|
|
| 2 |
18 UoC |
6 UoC |
|
|
| 2 |
12 UoC |
12 UoC |
|
|
| 3 |
12 UoC |
6 UoC |
6 UoC |
|
| 3 |
6 UoC |
12 UoC |
6 UoC |
|
| 3 |
6 UoC |
6 UoC |
12 UoC |
|
| 4 |
6 UoC |
6 UoC |
6 UoC |
6 UoC |
Things to take into account:
- Your other academic commitments. You are limited to a total of 18 UoC per term (without special exemption)
- Your other personal commitments. For example, if you are working full time, it won’t be possible to enroll in 18 UoC.
- The supervisor’s preference. The supervisor may need to complete the work quickly, and so prefer a two-term enrolment or there may be pragmatic things (like data access time frames) that make it preferable to run over 3-4 terms.
Examples of published manuscripts
Some students and supervisors go on to publish their research following the dissertation process. Click on the boxes below to explore examples.
Sharma A, Hanly M, Bhonagiri D. The impact of cultural and linguistic diversity on sepsis outcomes in patients admitted to ICUs: A multicenter, retrospective cohort study. Critical Care Medicine. 2026 Apr 13:10-97. doi.org/10.1097/ccm.0000000000007126
Objectives:
This study aims to investigate the effect of culturally and linguistically diverse (CaLD) status on in-hospital mortality in patients admitted to ICUs with sepsis. We hypothesize that diverse cultural and ethnic backgrounds, combined with limited English proficiency, might contribute to increased mortality in these patients.
Design:
Multicenter, retrospective cohort study.
Setting:
Adult ICUs with in South Western Sydney Local Health District (SWSLHD), New South Wales, Australia.
PATIENTS:
All adult patients 18 years or older, admitted to ICUs within the SWSLHD with a diagnosis of sepsis between January 1, 2012, and December 31, 2022.
MEASUREMENTS AND MAIN RESULTS:
Primary outcome was in-hospital mortality. ICU and hospital length of stays (LOSs) and readmission within 90 days to the ICU were our secondary outcomes. To isolate the effect of CaLD status on outcomes, matching was used to balance background covariates between the CaLD and non-CaLD groups. The average marginal effect of CaLD status on in-hospital mortality was then estimated using the matched data. In the analysis, 5971 sepsis-coded admissions were included, of which 2792 (46.75%) were from patients with CaLD backgrounds. Sixteen percent (435/2792) of the CaLD patients died in hospital compared with 17% (532/3179) deaths in the non-CaLD group. In the adjusted analysis on the matched data, hospital mortality was 2.2 percentage points lower (risk difference [RD], –0.022; 95% CI, 0.044 to –0.0005; p = 0.05) in the CaLD group compared with the non-CaLD group, corresponding to a 12.4% (risk ratio, 0.876; 95% CI, 0.763–0.989; p < 0.001) reduction in relative risk. ICU LOS was shorter for the CaLD patients by 0.53 days (12.72hr) (95% CI, –0.836 to –0.226; p < 0.001) compared with the non-CaLD group.
Conclusions:
Contrary to our hypothesis, in-hospital mortality after ICU admission with sepsis was lower in patients belonging to CaLD backgrounds. This effect was largely driven by patients from North African/Middle Eastern backgrounds, the largest CaLD subgroup.
Hao BS, Quiroz JC, Olver IN, Vajdic CM. The Association Between Alignment to the Breast Cancer Optimal Care Pathways and Patient Survival in Victoria, Australia, 2012–2019: A Retrospective Population‐Based Cohort Study. Medical Journal of Australia. 2026 Mar;224(3):e70162. doi.org/10.5694/mja2.70162
Objectives
To quantify the association between care alignment to the treatment step of the Cancer Council Victoria and Department of Health Victoria Optimal care pathways for people with breast cancer (OCP) (second edition) and survival.
Design
Retrospective population-based cohort study using the Victorian Cancer Registry and linked administrative health datasets.
Setting, Participants
Adult women diagnosed with invasive, unilateral breast cancer from 1 July 2012 to 31 December 2019 in Victoria, Australia.
Main Outcome Measures
Breast cancer-specific and overall survival for women whose care did or did not align to the treatment step of the OCP expressed as adjusted hazard ratios. Interaction between OCP alignment and cancer stage at diagnosis was also assessed.
Results
Of 29,591 eligible women, 17,152 (58.0%) were fully aligned, 7086 (23.9%) were partially aligned and 5353 (18.1%) were not aligned to the treatment step of the breast cancer OCP. Median follow-up was 1481 days (interquartile range, 850–2210 days). Adjusting for measured sociodemographic and clinical factors, OCP treatment alignment was associated with 23% (95% confidence interval [CI], 14%–31%) and 34% (95% CI, 29%–40%) lower risk of death from breast cancer and all causes, respectively, compared with non-alignment. By cancer stage, OCP alignment was significantly associated with 40% (95% CI, 27%–50%) and 30% (95% CI, 14%–42%) lower risk of breast cancer death for Stage II and III cancers, respectively, and 39% (95% CI, 26%–49%), 49% (95% CI, 42%–55%) and 33% (21%–44%) lower risk of all-cause death for Stage I, II and III cancers, respectively.
Conclusions
Risk of death was lower for women with breast cancer whose treatment aligned to the OCP compared with women whose treatment did not align. Our findings support the promotion and implementation of the breast cancer OCP.
Luo J, Woodward M, Ferreira ML, Harris K. Sex differences in the experience of pain in the UK Biobank cohort study. British Journal of Pain. 2026 Jan 22:20494637261418196. doi.org/10.1177/20494637261418196
Objective
Chronic pain has been shown to be more prevalent among women than men. However, each person’s experience of pain is shaped by a complex interplay of biological, psychological and social factors. The aim of this study was to summarise a comprehensive ‘experience of pain’ questionnaire in the UK Biobank and identify differences in the experience of chronic pain between females and males.
Methods
This was an exploratory analysis of an online self-assessment questionnaire consisting of 128 questions related to UK Biobank participant’s experience of pain that was administered in 2019. Data were summarised by sex, and chi-squared and t-tests were used to determine whether there were statistically significant differences between females and males.
Results
About one-third of UK Biobank participants (167,183, 57% female) responded to the questionnaire. More females than males reported suffering from chronic pain (60.0% vs 51.5%). There was female predominance in 11 out of 14 medical conditions, particularly in osteoarthritis (35.6% females vs 24.5% males), migraine (25.0% vs 12.3%) and fibromyalgia (2.7% vs 0.7%). Female participants tended to report pain of greater severity and longer duration that more profoundly impairs their everyday functioning when compared to their male counterparts.
Conclusion
A significant strength of our study is the large sample size, and the high detail of information captured about pain phenotypes, in which we found sex differences in chronic pain persist. We recommend future pain surveys collect sex-based pain conditions to enable better recognition of why sex differences in pain persist.
Muniratul MM, Bliuc D, Mather K, Ivers R, Sachdev P, Broadaty H, Dai-Keller Z. The association between fermented dairy food intake and incident depression and dementia in older Australians. Proceedings of the Nutrition Society. 2025 Apr;84(OCE1):E104. doi.org/10.1017/S0029665125001144
Abstract
Depression and dementia represent significant public health issues, affecting approximately 1 in 10 and 1 in 12 older Australians, respectively. While current pharmacological treatments are effective in relieving symptoms, they often entail undesirable adverse effects, including gastrointestinal issues and bradycardia(1,2). This highlights the need for primary preventative measures, including food- and nutrition-based approaches. Chronic brain inflammation is believed to interfere with the gut–brain axis(3). Consumption of fermented dairy products rich in beneficial gut microbes may attenuate this inflammation and offer protective health benefits. This study aimed to examine whether fermented dairy intake could mitigate the risk of incident depression and dementia. Utilising data from the Sydney Memory and Ageing Study I of 1037 participants 70–90 years, 816 participants (mean age: 76.7) were followed from 2005 until 2012 for incident depression, and 974 participants (mean age: 80.7) were followed up from 2005 until 2014 for incident dementia. Fermented dairy intake was assessed using the Dietary Questionnaire for Epidemiological Studies version 2 and categorised yoghurt and regular cheese into quartiles (Q) and low-fat cheese into consumers/non-consumers, with no consumption as the reference group. Depression diagnoses were assessed via self-reported physician-diagnosed history, medication use, service utilisation, and heavy alcohol use. Dementia diagnoses followed the criteria in the fourth edition of the Diagnostic and Statistical Manual of Mental Disorders. Cox proportional hazards models examined the associations between fermented dairy intake and the risk of incident depression/dementia. Additionally, linear regression models were applied to assess for depressive symptoms score (measured by the Geriatric Depression Scale-15) and psychological distress score (measured by the Kessler Psychological Distress Scale-10). All models were adjusted for sociodemographic, lifestyle factors, and medical histories. Over a median follow-up of 3.9 and 5.8 years, 120 incident depression and 100 incident dementia cases occurred, respectively. Those who consumed high yoghurt (Q4: 145.8–437.4 g/day) and low-fat cheese (consumers: 0.4–103.1g/day) intakes were associated with a lower risk of incident depression, both compared to non-consumers (yoghurt: adj.HR: 0.38, 95% CI: 0.19–0.77; low-fat cheese: adj.HR: 0.50; 95% CI: 0.29–0.86). They were also associated with lower depressive symptom scores (yoghurt: adj.β = −0.46; 95% CI: −0.84, −0.07; low-fat cheese: adj.β = −0.42; 95% CI: −0.73, −0.11). However, those who consumed a higher intake of regular cheese (Q4: 14.7–86.1 g/day) had an elevated risk of incident depression (adj.HR: 1.88; 95% CI: 1.02, 3.47), and those in Q2 (0.1–7.2 g/day) had significantly higher depressive symptom scores (adj.β = 0.42; 95% CI: 0.05, 0.78). No significant findings were found for psychological distress scores or incident dementia. Our findings of a cohort of older Australians suggest that higher yoghurt and low-fat cheese intakes may reduce incident depression and depressive symptoms, while a higher intake of regular cheese may increase these risks.
Farey JE, Li A, Adie S, Smith PN, Lujic S, Harris IA. The cumulative incidence of dislocation and revision surgery following total hip arthroplasty for hip fracture in New South Wales: a data linkage study. The Bone & Joint Journal. 2025 Oct 1;107(10):1004-10. doi.org/10.1302/0301-620X.107B10.BJJ-2024-1637.R1
Aims
Dislocation is a common problem after total hip arthroplasty (THA) for hip fracture. This study aimed to assess the one-year cumulative incidence of dislocation, and identify associated risk factors.
Methods
An observational cohort study was conducted using data from the Australian Orthopaedic Association National Joint Replacement Registry linked with the New South Wales Admitted Patient Data Collection. Patients aged over 18 years who underwent THA for fracture neck of femur between 1 July 2010 and 31 December 2018 in New South Wales were included. Dislocations and revision surgeries were identified via linked datasets. Multivariable logistic regression evaluated demographic, surgical, and implant-related risk factors for dislocation. Subgroup analysis considered surgical approach and BMI.
Results
Among 4,632 patients, the one-year dislocation incidence was 4.8% (95% CI 4.2 to 5.5), with 79% occurring within 90 days. Revision for dislocation occurred in 1.1% of cases (95% CI 0.80 to 1.4). Compared with dual-mobility acetabular components, conventional bearings ≤ 32 mm (odds ratio (OR) 1.64 (95% CI 0.93 to 2.90); p = 0.087) and > 32 mm (OR 1.33 (95% CI 0.75 to 2.37); p = 0.332) showed no significant difference in dislocation risk. In a subgroup of 2,532 patients, the anterior approach significantly reduced dislocation risk (OR 0.28 (95% CI 0.12 to 0.67); p = 0.004), whereas the lateral approach did not (OR 0.75 (95% CI 0.48 to 1.17); p = 0.202) compared to the posterior approach. Adjusting for surgical approach, ≤ 32 mm bearings were associated with a higher dislocation risk than dual-mobility components (OR 1.97 (95% CI 1.06 to 3.66); p = 0.031); > 32 mm bearings were not significantly different (OR 1.68 (95% CI 0.89 to 3.15); p = 0.110).
Conclusion
One in 20 patients undergoing THA for fracture will experience dislocation within a year, though most will not require revision. Dual-mobility components may be protective against dislocation compared with smaller-diameter femoral head sizes.
Liu Q, Han Y, Shen L, Du J, Tania MH. Adaptive enhancement of shoulder x-ray images using tissue attenuation and type-II fuzzy sets. PlOS One. 2025 Feb 6;20(2):e0316585. doi.org/10.1371/journal.pone.0316585
Abstract
Shoulder X-ray images typically have low contrast and high noise levels, making it challenging to distinguish and identify subtle anatomical structures. While existing image enhancement techniques are effective in improving contrast, they often overlook the enhancement of sharpness, especially when amplifying blurring and noise. These techniques may improve detail contrast but fail to maintain overall image clarity and the distinction between the target and the background. To address these issues, we propose a novel image enhancement method aimed at simultaneously improving both the contrast and sharpness of shoulder X-ray images. The method integrates automatic tissue attenuation techniques, which enhance the image contrast by removing non-essential tissue components while preserving important tissues and bones. Additionally, we apply an improved Type-II fuzzy set algorithm to further optimize image sharpness. By simultaneously enhancing contrast and sharpness, the method significantly improves image quality and detail distinguishability. When tested on certain images from the MURA dataset, the proposed method achieved the best or second-best results, outperforming five no-reference image quality assessment metrics. In comparative studies, the method demonstrated significant performance advantages over 10 contemporary X-ray image enhancement algorithms and was validated through ablation experiments to confirm the effectiveness of each module.
Liu Q, Zhou T, Cheng C, Ma J, Hoque Tania M. Hybrid generative adversarial network based on frequency and spatial domain for histopathological image synthesis. BMC Bioinformatics. 2025 Jan 27;26(1):29. doi.org/10.1186/s12859-025-06057-9
Background
Due to the complexity and cost of preparing histopathological slides, deep learning-based methods have been developed to generate high-quality histological images. However, existing approaches primarily focus on spatial domain information, neglecting the periodic information in the frequency domain and the complementary relationship between the two domains. In this paper, we proposed a generative adversarial network that employs a cross-attention mechanism to extract and fuse features across spatial and frequency domains. The method optimizes frequency domain features using spatial domain guidance and refines spatial features with frequency domain information, preserving key details while eliminating redundancy to generate high-quality histological images.
Results
Our model incorporates a variable-window mixed attention module to dynamically adjust attention window sizes, capturing both local details and global context. A spectral filtering module enhances the extraction of repetitive textures and periodic structures, while a cross-attention fusion module dynamically weights features from both domains, focusing on the most critical information to produce realistic and detailed images.
Conclusions
The proposed method achieves efficient spatial-frequency domain fusion, significantly improving image generation quality. Experiments on the Patch Camelyon dataset show superior performance over eight state-of-the-art models across five metrics. This approach advances automated histopathological image generation with potential for clinical applications.
Bi B, Liu L, Perez-Concha O. Adapting large language models for automated summarisation of electronic medical records in clinical coding. In Health. Innovation. Community: It Starts With Us: Papers from the 28th Australian Digital Health and Health Informatics Conference (HIC 2024) doi.org/10.3233/SHTI240886
Abstract
Early identification of vulnerable children to protect them from harm and support them in achieving their long-term potential is a community priority. This is particularly important in the Northern Territory (NT) of Australia, where Aboriginal children are about 40% of all children, and for whom the trauma and disadvantage experienced by Aboriginal Australians has ongoing intergenerational impacts. Given that shared social determinants influence child outcomes across the domains of health, education and welfare, there is growing interest in collaborative interventions that simultaneously respond to outcomes in all domains. There is increasing recognition that many children receive services from multiple NT government agencies, however there is limited understanding of the pattern and scale of overlap of these services. In this paper, NT health, education, child protection and perinatal datasets have been linked for the first time. The records of 8,267 children born in the NT in 2006–2009 were analysed using a person-centred analytic approach. Unsupervised machine learning techniques were used to discover clusters of NT children who experience different patterns of risk. Modelling revealed four or five distinct clusters including a cluster of children who are predominantly ill and experience some neglect, a cluster who predominantly experience abuse and a cluster who predominantly experience neglect. These three, high risk clusters all have low school attendance and together comprise 10–15% of the population. There is a large group of thriving children, with low health needs, high school attendance and low CPS contact. Finally, an unexpected cluster is a modestly sized group of non-attendees, mostly Aboriginal children, who have low school attendance but are otherwise thriving. The high risk groups experience vulnerability in all three domains of health, education and child protection, supporting the need for a flexible, rather than strictly differentiated response. Interagency cooperation would be valuable to provide a suitably collective and coordinated response for the most vulnerable children.Abstract
Bi B, Liu L, Lujic S, Jorm L, Perez-Concha O. Harmonising the clinical melody: tuning large language models for hospital course summarisation in clinical coding. arXiv preprint arXiv:2409.14638. 2024 Sep 23. doi.org/10.48550/arXiv.2409.14638
Abstract
The increasing volume and complexity of clinical documentation in Electronic Medical Records (EMR) systems pose significant challenges for clinical coders, who must mentally process and summarise vast amounts of clinical text to extract essential information needed for coding tasks. While large language models (LLMs) have been successfully applied to shorter summarisation tasks in recent years, the challenge of summarising a hospital course remains an open area for further research and development. In this study, we adapted three pretrained LLMs (Llama 3, BioMistral, Mistral Instruct v0.1) for the hospital course summarisation task, using Quantized Low-Rank Adaptation (QLoRA) fine-tuning. We created a free-text clinical dataset from MIMIC III data by concatenating various clinical notes as the input clinical text, paired with ground truth ‘Brief Hospital Course’ sections extracted from the discharge summaries for model training. The fine-tuned models were evaluated using BERTScore and ROUGE metrics to assess the effectiveness of clinical domain fine-tuning. Additionally, we validated their practical utility using a novel hospital course summary assessment metric specifically tailored for clinical coding. Our findings indicate that fine-tuning pre-trained LLMs for the clinical domain can significantly enhance their performance in hospital course summarisation and suggest their potential as assistive tools for clinical coding. Future work should focus on refining data curation methods to create higher quality clinical datasets tailored for hospital course summary tasks and adapting more advanced open-source LLMs comparable to proprietary models to further advance this research.
Shala DR, Sheppard‐Law S. Measuring text similarity and its associated factors in electronic nursing progress notes: A retrospective review. Journal of Clinical Nursing. 2023 Jun;32(11-12):2733-41. doi.org/10.1111/jocn.16374
Aims and objectives
To measure text similarity in electronic nursing progress notes and determine factors associated with text similarity.
Background
Electronic clinical notes with redundant information masks clinically relevant information, increases clinicians’ cognitive burden and undermines patient safety.
Design
Retrospective review of electronic medical record nursing progress notes.
Methods
The study was conducted between November 2018 and February 2019 in two Australian Paediatric Intensive Care Units. De-identified, randomly selected inpatient data were extracted from the network’s database. Manually classified shift summary progress notes for each admission were sequenced from admission to discharge. Text similarity was calculated for consecutive pairs of nursing progress notes. Linear regression was undertaken to determine the association between the similarity scores and variables of interest: note word count, total number of notes and unit. The STROBE checklist was used for reporting.
Results
921 shift summary nursing progress notes were analysed. Similarity scores were widely distributed with a median of 10.37%. Only 17.2% (n = 144) of the notes have similarity scores above 20%. Of these, 5% (n = 47) were above 50% similar in comparison with a previously written note. Similarity above 50% was observed as early as the first note pair in the course of a patient’s admission. A significant difference was found between the similarity scores of Unit 1 and Unit 2. Hospital unit was the only variable of interest significantly associated with similarity scores.
Conclusion
Text similarity among electronic nursing progress notes in Australian Paediatric ICUs is minimal; however, notes with >50% similarity have been identified. Text analytics provides measurable data and insights about electronic clinical documentation to inform future nursing practice, research and eMR design.
Relevance to clinical practice
Findings have implications for nursing practice in the way that nursing staff are educated to maintain data quality, professional accountability and effective communication in electronic documentation and to avoid unnecessary repetition of text.
Nicholson C, Hanly M, Celermajer DS. An interactive geographic information system to inform optimal locations for healthcare services. PLOS Digital Health. 2023 May 8;2(5):e0000253. doi.org/10.1371/journal.pdig.0000253
Abstract
Large health datasets can provide evidence for the equitable allocation of healthcare resources and access to care. Geographic information systems (GIS) can help to present this data in a useful way, aiding in health service delivery. An interactive GIS was developed for the adult congenital heart disease service (ACHD) in New South Wales, Australia to demonstrate its feasibility for health service planning. Datasets describing geographic boundaries, area-level demographics, hospital driving times, and the current ACHD patient population were collected, linked, and displayed in an interactive clinic planning tool. The current ACHD service locations were mapped, and tools to compare current and potential locations were provided. Three locations for new clinics in rural areas were selected to demonstrate the application. Introducing new clinics changed the number of rural patients within a 1-hour drive of their nearest clinic from 44·38% to 55.07% (79 patients) and reduced the average driving time from rural areas to the nearest clinic from 2·4 hours to 1·8 hours. The longest driving time was changed from 10·9 hours to 8·9 hours. A de-identified public version of the GIS clinic planning tool is deployed at https://cbdrh.shinyapps.io/ACHD_Dashboard/. This application demonstrates how a freely available and interactive GIS can be used to aid in health service planning. In the context of ACHD, GIS research has shown that adherence to best practice care is impacted by patients’ accessibility to specialist services. This project builds on this research by providing opensource tools to build more accessible healthcare services.
Roper L, He VY, Perez-Concha O, Guthridge S. Complex early childhood experiences: characteristics of Northern Territory children across health, education and child protection data. PLOS One. 2023 Jan 19;18(1):e0280648. doi.org/10.1371/journal.pone.0280648
Abstract
Early identification of vulnerable children to protect them from harm and support them in achieving their long-term potential is a community priority. This is particularly important in the Northern Territory (NT) of Australia, where Aboriginal children are about 40% of all children, and for whom the trauma and disadvantage experienced by Aboriginal Australians has ongoing intergenerational impacts. Given that shared social determinants influence child outcomes across the domains of health, education and welfare, there is growing interest in collaborative interventions that simultaneously respond to outcomes in all domains. There is increasing recognition that many children receive services from multiple NT government agencies, however there is limited understanding of the pattern and scale of overlap of these services. In this paper, NT health, education, child protection and perinatal datasets have been linked for the first time. The records of 8,267 children born in the NT in 2006–2009 were analysed using a person-centred analytic approach. Unsupervised machine learning techniques were used to discover clusters of NT children who experience different patterns of risk. Modelling revealed four or five distinct clusters including a cluster of children who are predominantly ill and experience some neglect, a cluster who predominantly experience abuse and a cluster who predominantly experience neglect. These three, high risk clusters all have low school attendance and together comprise 10–15% of the population. There is a large group of thriving children, with low health needs, high school attendance and low CPS contact. Finally, an unexpected cluster is a modestly sized group of non-attendees, mostly Aboriginal children, who have low school attendance but are otherwise thriving. The high risk groups experience vulnerability in all three domains of health, education and child protection, supporting the need for a flexible, rather than strictly differentiated response. Interagency cooperation would be valuable to provide a suitably collective and coordinated response for the most vulnerable children.
Frequently Asked Questions
The MSc Health Data Science offers a choice between a 24 Units of Credit workplace/internship research dissertation (full-time or part-time options) or a 6 Units of Credit capstone project plus 3 x 6 Units of Credit electives (from a selection of over 20 courses – see Handbook). The choice of a pathway depends on your desire to work independently on a larger project (dissertation) or a more pre-defined research project (capstone). Students wishing to progress towards a PhD are encouraged to enroll in a dissertation pathway.
Gaining a place in a dissertation is competitive and at the discretion of both the course convenor(s) and the supervisor(s) offering a particular dissertation project.
| Description |
You conduct an independent research project under the supervision of a UNSW academic |
You apply the health data science pipeline to a pre-defined research project |
| Research question |
Proposed by or co-developed with a supervisor; unique for each student |
Provided to you by the course convenor; the same for every student |
| Opportunity to publish a research paper |
Yes, there is potential |
Limited opportunity |
| Assessment |
There are three components: A project protocol (20%), a presentation (20%) and the final submitted output—usually a draft research manuscript, research report, interactive app or data visualisation (60%) |
There are three components: Data access (5%), Project plan (15%), Final report (80%) |
| Total Units of Credit |
24 |
6 |
| Time to complete |
2 Terms full time or 3-4 Terms part time |
1 Term |
Towards the end of every term the conveners open up an expression of interest process for students interested in the dissertation course. Keep an eye out for this. Broadly, there are three ways to arrange a project:
- A workplace project Students propose a project with a supervisor from their workplace. This is contingent on the student having an established relationship and support from their workplace.
- Approaching a potential supervisor Students approach a potential supervisor at UNSW and propose a project or co-develop a project idea with the supervisor. This works best for students who have a very strong interest in a particular area and can demonstrate that to the potential supervisor.
- Competitive application Students apply for projects that are proposed by supervisors. The proposed projects will be released towards the end of each term to students who have completed the expression of interest process. Securing a project is a competitive process: students indicate projects they are interested in and the supervisor selects the student they want to work on their project. This is the most common way that students are placed in projects but it does involve a level of competition and uncertainty.
There are four deliverables, three of which are assessed.
| Student-supervisor agreement |
First term of study: Friday of Week 3 |
0% |
| Project protocol |
First Term of Study: Sunday of Week 6 for students enrolled in 12 or 18 UoC and Sunday of Week 10 for students enrolled in 6 UoC |
20% |
| Oral presentation |
Final Term of Study: Week 9 |
20% |
| Final Output |
Final Term of Study: Week 10 |
60% |
All projects must use health data and data science analytics and should reflect the type of projects that health data scientists might work on in real-life workplace settings. For this reason, project types are likely to vary between students.
The project output might be a traditional type of research project involving a series of analyses of a large dataset, it might be focused on generating a health data science tool, or it might be a hybrid of these two main types. In the first case the output may be a thesis or manuscript draft. In the second case, the output will be the tool itself and a supporting technical report. The tool must be in some way accessible to the dissertation examiners, for example a dashboard deployed to a public URL.
If there is limited access to the dissertation output, for example a dashboard deployed in a Trusted Research Environment, the supervisor(s) should notify the Course Convenor as soon as possible, as this will restrict the pool of potential examiners.
By supervising a Health Data Science dissertation you commit to the following:
- Regular meetings with the student (typically ranging between 1 hour per week for full time students to 30 minutes per fortnight for part-time students).
- Proving marks and feedback for the study protocol and final output.
- Nominating and contacting an external examiner to examine the final output component of the dissertation (this should be someone outside of your immediate research team)
Examples of past topics
- Investigating the role of repeat expansions in the causation of spontaneous coronary artery dissection
- Measuring adherance to multiple medicines: a case study among people on combination antiretroviral therapy in Australia
- Secondary analysis of MedicineInsight dataset to explore the demographic and health profiles of the people with severe mental illness, their health-seeking behaviour, and management in primary care settings
- Predicting negative clinical outcomes following transfer from acute to subacute care in a geriatric population
- Identifying timeliness and factors associate with the completion of an electronic adult admission assessment in Sydney hospitals
- Disambiguate, Model, Recommend & Visualise: A proof of concept toolkit to explore research collaboration
- Profiling chronic disease and health service utilisation of elderly hospitalised NSW patients
- How multiple failures at the practical driving test and hazard perception test influences crash risk over a 13-year period: the DRIVE study
- Developing an interactive geographic information system for Adult Congenital Heart Disease service planning in rural NSW
- Bayesian Analysis in Drug and Alcohol Treatment: a case study
- Sex differences in Health Outcomes following Myocardial Infarction in patients managed by Emergency Medical Service
- TreatmentEstimatoR: a dashboard for estimating treatment effects from observational health data
- Use of co-medications and potential drug-drug interactions amongst people living with HIV on antiretroviral therapy in Australia
- The application of Causal Inference methods to evaluate Multiple Programs delivered in the NSW Child Protection Service System
- Visualising and monitoring global COVID-19 vaccine inequity
- Predicting outcomes in total knee replacement surgery
- Projecting future demand for health services and effectively communicating the results
- Complex early childhood experiences: characteristics of Northern Territory children across education, health and child protection data
- Predicting mortality after discharge from hospital in cardiac patients using machine learning on electronic medical records (EMR).
- Natural language processing of clinical notes to extract details of treatment with orally delivered chemotherapy and targeted therapies
- Forecasting All-Cause Mortality: Leveraging Causes-of-Death Data through Neural Networks
- Differences in revascularisation rates between publicly and privately funded patients admitted to NSW public hospitals for acute myocardial infarction
- Unpacking the propensity score matching paradox
- Understanding the environmental health of food packaging in Australia
- Sex differences in the prevalence and impact of frailty in heart failure patients
- Changes in cardiovascular disease risk over time among people living with HIV
- Systematic data quality assessment of general practitioner collected smoking information on smoking associated disease risks using MedicineInsight data
- Understanding trends in illicit drug use and harms across Australia: Developing a data visualisation tool to inform policy
- Automated dashboard to support the management of pain, malnutrition, depression and anxiety in the care of people with cancer *Forecasting all-cause mortality: leveraging Cause-of-Death data through neural networks
- Sex differences in pain and multimorbidity in the UKBiobank
- Examining the relationship between fruits and vegetables intake and DNA methylation in Australian older population
- The epidemiology of falls in children and adolescents in NSW 2001-2019: findings from linked administrative data
- Detecting semantic relationships for clinical concepts contained in free-text clinical notes and other documents in electronic medical record systems
- African regional GIS based analysis for spatio-temporal trends of zoonotic diseases using artificial intelligence data
- Prompt-engineering A Large Language Model for Data Extraction from New South Wales Coronial Findings and Recommendations
- Classification of Causes of Death in the Australian and New Zealand Neonatal Network Clinical Quality Registry
- Comparative Evaluation of Comorbidity and Frailty Indices for Predicting In-Hospital and One-Year Mortality
- Identifying Intellectual Disability in Discharge Summaries Using Large Language Models and Prompt Engineering
- Automated clinical coding using large language models
- Automated concept extraction from electronic medical records using prompt engineering
- Adapting large language models for concept extraction from electronic medical records
- Dashboard for the Australian and New Zealand Neonatal Network (ANZNN) data registry
- Automated methods for summarization of electronic medical records for clinical coding
- Complex early childhood experiences: characteristics of Northern Territory children across education, health and child protection data
- The Application of Causal Inference Methods to Evaluate Multiple Programs Delivered in the NSW Child Protection Service System
- Predicting negative clinical outcomes following transfer from acute to subacute care in a geriatric population
- Exploring the use of interactive visualisation tools to improve amphetamine treatment in New South Wales