95 lines
4.7 KiB
TeX
95 lines
4.7 KiB
TeX
%! Author = alex
|
|
%! Date = 3/6/25
|
|
|
|
|
|
\section{Results}\label{sec:results}
|
|
|
|
We begin by comparing the overall performance of all trained models and baselines using four key evaluation metrics:
|
|
mean absolute error (MAE), mean squared error (MSE), and the coefficient of determination (\(R^2\)) for the regression target,
|
|
as well as the metrics for their respective sub-intervals.
|
|
Figure~\ref{fig:results_performance_overview_by_model_type} provides an overview of these metrics across all model types.
|
|
|
|
Overall, transformer-based models consistently outperformed LSTM variants and baseline methods across most evaluation criteria.
|
|
Among the baselines, [e.g., "the rule-based method"] showed the weakest performance,
|
|
while the [e.g., "windowed logistic regression"] performed competitively in certain contexts.
|
|
Differences across models were most pronounced in MSE and \(R^2\),
|
|
indicating that advanced architectures better captured higher-order dynamics and reduced large prediction errors.
|
|
|
|
The previous section detailed the design and implementation of our fertility prediction pipeline,
|
|
including data preprocessing, feature engineering, input encoding, and the development of several deep learning architectures.
|
|
We now present the results of our evaluation, focusing on the predictive accuracy of the proposed models across different temporal resolutions,
|
|
cycle types and use cases.
|
|
Model performance is assessed using both overall metrics and biologically targeted subintervals,
|
|
allowing for a nuanced comparison of approaches and their practical relevance to real-time fertility forecasting.
|
|
Additionally, model performance is compared to the three baseline models introduced.
|
|
|
|
% provide information about the training behaviour and statistic of the different models??
|
|
|
|
\subsection{Overall Model Performance Across Architectures}\label{subsec:overall_model_performance_across_architectures}
|
|
|
|
\begin{figure}
|
|
\centering
|
|
\includegraphics[width=0.9\textwidth]{resources/figures/results/model_performance_overview}
|
|
\caption{Overview of the performances of all model types including the baselines on 4 selected performance metrics.}
|
|
\label{fig:results_performance_overview_by_model_type}
|
|
\end{figure}
|
|
|
|
\subsection{Fertility Probability Prediction Accuracy}\label{subsec:fertility_probability_precition_accuracy}
|
|
|
|
\begin{landscape}
|
|
\begin{table}[ht]
|
|
\centering
|
|
\caption{Model comparison for fertility prediction using MAE and MSE}
|
|
\begin{adjustbox}{max width=\linewidth}
|
|
\begin{tabular}{lllcccccc}
|
|
\toprule
|
|
\textbf{Model} & \textbf{Window} & \textbf{Daily} &
|
|
\textbf{MAE$_{fert}$} & \textbf{MAE$_{fert,during}$} & \textbf{MAE$_{fert,non}$} &
|
|
\textbf{MSE$_{fert}$} & \textbf{MSE$_{fert,during}$} & \textbf{MSE$_{fert,non}$} \\
|
|
\midrule
|
|
ModelA & 7 & Yes & 0.67 & 0.59 & 0.73 & 0.89 & 0.82 & 0.95 \\
|
|
ModelB & 14 & No & 0.65 & 0.58 & 0.71 & 0.87 & 0.80 & 0.93 \\
|
|
% More rows...
|
|
\bottomrule
|
|
\end{tabular}
|
|
\end{adjustbox}
|
|
\label{tab:fertility_comparison}
|
|
\end{table}
|
|
\end{landscape}
|
|
|
|
|
|
% show why I selected the individual input configs for model config training
|
|
% selected by best mse fertility, use 2nd best, as it provides basically the same performance, but more input data for more complex model configs
|
|
|
|
\subsubsection{Fertility Probability Prediction}\label{subsubsec:fertility_probability_prediction}
|
|
|
|
\subsubsection{Impact of Input Resolution}\label{subsubsec:fert_impact_of_input_resolution}
|
|
|
|
\subsubsection{Impact of Input Window Length}\label{subsubsec:fert_impact_of_historical_context}
|
|
|
|
\subsubsection{Comparison with Baselines}\label{subsubsec:fert_comparison_with_baselines}
|
|
|
|
\subsection{Ovulation-Over Prediction}\label{subsubsec:ov_over_prediction}
|
|
|
|
\subsubsection{Performance around Ovulation}\label{subsubsec:ov_over_performance_around_ovulation}
|
|
|
|
\subsubsection{Impact of Input Resolution}\label{subsubsec:ov_over_impact_of_input_resolution}
|
|
|
|
\subsubsection{Impact of Input Window Length}\label{subsubsec:ov_over_impact_of_historical_context}
|
|
|
|
\subsubsection{Comparison with Baselines}\label{subsubsec:ov_over_comparison_with_baselines}
|
|
|
|
\subsection{Stratified Analysis}\label{subsec:stratified_analysis}
|
|
|
|
\subsubsection{Regular vs Irregular Cycles}\label{subsubsec:regular_vs_irregular_cycles}
|
|
|
|
\subsubsection{Influence of User History Depth}\label{subsubsec:influence_of_past_user_data}
|
|
|
|
\subsection{Use-Case Evaluation Results}\label{subsec:use_case_evaluation_results}
|
|
|
|
\subsubsection{Contraception Use-Case Results}\label{subsubsec:use_case_contraception_results}
|
|
|
|
\subsubsection{Pregnancy Use-Case Results}\label{subsubsec:use_case_pregnancy_results}
|
|
|
|
\subsection{Summary of Key Findings}\label{subsec:summary_of_key_findings}
|