Databricks-Machine-Learning-Associate Databricks Certified Machine Learning Associate Exam Questions and Answers

Questions 4

A data scientist is using MLflow to track their machine learning experiment. As a part of each of their MLflow runs, they are performing hyperparameter tuning. The data scientist would like to have one parent run for the tuning process with a child run for each unique combination of hyperparameter values. All parent and child runs are being manually started with mlflow.start_run.

Which of the following approaches can the data scientist use to accomplish this MLflow run organization?

Options:

Theycan turn on Databricks Autologging

Theycan specify nested=True when startingthe child run for each unique combination of hyperparameter values

Theycan start each child run inside the parentrun's indented code block usingmlflow.start runO

They can start each child run with the same experiment ID as the parent run

They can specify nested=True when starting the parent run for the tuningprocess

Buy Now

Questions 5

A data scientist wants to use Spark ML to one-hot encode the categorical features in their PySpark DataFramefeatures_df. A list of the names of the string columns is assigned to theinput_columnsvariable.

They have developed this code block to accomplish this task:

Databricks-Machine-Learning-Associate Question 5

The code block is returning an error.

Which of the following adjustments does the data scientist need to make to accomplish this task?

Options:

They need to specify the method parameter to the OneHotEncoder.

They need to remove the line with the fit operation.

They need to use Stringlndexer prior to one-hot encodinq the features.

They need to useVectorAssemblerprior to one-hot encoding the features.

Buy Now

Questions 6

A machine learning engineer has grown tired of needing to install the MLflow Python library on each of their clusters. They ask a senior machine learning engineer how their notebooks can load the MLflow library without installing it each time. The senior machine learning engineer suggests that they use Databricks Runtime for Machine Learning.

Which of the following approaches describes how the machine learning engineer can begin using Databricks Runtime for Machine Learning?

Options:

They can add a line enabling Databricks Runtime ML in their init script when creating their clusters.

They can check the Databricks Runtime ML box when creating their clusters.

They can select a Databricks Runtime ML version from the Databricks Runtime Version dropdown when creating their clusters.

They can set the runtime-version variable in their Spark session to “ml”.

Buy Now

Questions 7

A data scientist has created two linear regression models. The first model uses price as a label variable and the second model uses log(price) as a label variable. When evaluating the RMSE of each model bycomparing the label predictions to the actual price values, the data scientist notices that the RMSE for the second model is much larger than the RMSE of the first model.

Which of the following possible explanations for this difference is invalid?

Options:

The second model is much more accurate than the first model

The data scientist failed to exponentiate the predictions in the second model prior tocomputingthe RMSE

The datascientist failed to take the logof the predictions in the first model prior to computingthe RMSE

The first model is much more accurate than the second model

The RMSE is an invalid evaluation metric for regression problems

Buy Now

Questions 8

A data scientist has been given an incomplete notebook from the data engineering team. The notebook uses a Spark DataFrame spark_df on which the data scientist needs to perform further feature engineering. Unfortunately, the data scientist has not yet learned the PySpark DataFrame API.

Which of the following blocks of code can the data scientist run to be able to use the pandas API on Spark?

Options:

import pyspark.pandas as ps

df = ps.DataFrame(spark_df)

import pyspark.pandas as ps

df = ps.to_pandas(spark_df)

spark_df.to_pandas()

import pandas as pd

df = pd.DataFrame(spark_df)

Buy Now

Questions 9

Which of the following describes the relationship between native Spark DataFrames and pandas API on Spark DataFrames?

Options:

pandas API on Spark DataFrames are single-node versions of Spark DataFrames with additional metadata

pandas API on Spark DataFrames are more performant than Spark DataFrames

pandas API on Spark DataFrames are made up of Spark DataFrames and additional metadata

pandas API on Spark DataFrames are less mutable versions of Spark DataFrames

Buy Now

Questions 10

Which of the following hyperparameter optimization methods automatically makes informed selections of hyperparameter values based on previous trials for each iterative model evaluation?

Options:

Random Search

Halving Random Search

Tree of Parzen Estimators

Grid Search

Buy Now

Questions 11

Which of the following approaches can be used to view the notebook that was run to create an MLflow run?

Options:

Open the MLmodel artifact in the MLflow run paqe

Click the "Models" link in the row corresponding to the run in the MLflow experiment paqe

Click the "Source" link in the row corresponding to the run in the MLflow experiment page

Click the "Start Time" link in the row corresponding to the run in the MLflow experiment page

Buy Now

Questions 12

The implementation of linear regression in Spark ML first attempts to solve the linear regression problem using matrix decomposition, but this method does not scale well to large datasets with a large number of variables.

Which of the following approaches does Spark ML use to distribute the training of a linear regression model for large data?

Options:

Logistic regression

Spark ML cannot distribute linear regression training

Iterative optimization

Least-squares method

Singular value decomposition

Buy Now

Questions 13

A team is developing guidelines on when to use various evaluation metrics for classification problems. The team needs to provide input on when to use the F1 score over accuracy.

Databricks-Machine-Learning-Associate Question 13

Which of the following suggestions should the team include in their guidelines?

Options:

The F1 score should be utilized over accuracy when the number of actual positive cases is identical to the number of actual negative cases.

The F1 score should be utilized over accuracy when there are greater than two classes in the target variable.

The F1 score should be utilized over accuracy when there is significant imbalance between positive and negative classes and avoiding false negatives is a priority.

The F1 score should be utilized over accuracy when identifying true positives and true negatives are equally important to the business problem.

Buy Now

Questions 14

A data scientist wants to parallelize the training of trees in a gradient boosted tree to speed up the training process. A colleague suggests that parallelizing a boosted tree algorithm can be difficult.

Which of the following describes why?

Options:

Gradient boosting is not a linear algebra-based algorithm which is required for parallelization

Gradient boosting requires access to all data at once which cannot happen during parallelization.

Gradient boosting calculates gradients in evaluation metrics using all cores which prevents parallelization.

Gradient boosting is an iterative algorithm that requires information from the previous iteration to perform the next step.

Buy Now

Answer:

Explanation:

Gradient boosting is fundamentally an iterative algorithm where each new tree is built based on the errors of the previous ones. This sequential dependency makes it difficult to parallelize the training of trees in gradient boosting, as each step relies on the results from the preceding step. Parallelization in this context would undermine the core methodology of the algorithm, which depends on sequentially improving the model'sperformance with each iteration.References:

Machine Learning Algorithms (Challenges with Parallelizing Gradient Boosting).

Gradient boosting is an ensemble learning technique that builds models in a sequential manner. Each new model corrects the errors made by the previous ones. This sequential dependency means that each iteration requires the results of the previous iteration to make corrections. Here is a step-by-step explanation of why this makes parallelization challenging:

Sequential Nature: Gradient boosting builds one tree at a time. Each tree is trained to correct the residual errors of the previous trees. This requires the model to complete one iteration before starting the next.
Dependence on Previous Iterations: The gradient calculation at each step depends on the predictions made by the previous models. Therefore, the model must wait until the previous tree has been fully trained and evaluated before starting to train the next tree.
Difficulty in Parallelization: Because of this dependency, it is challenging to parallelize the training process. Unlike algorithms that process data independently in each step (e.g., random forests), gradient boosting cannot easily distribute the work across multiple processors or cores for simultaneous execution.

This iterative and dependent nature of the gradient boosting process makes it difficult to parallelize effectively.

References

Gradient Boosting Machine Learning Algorithm
Understanding Gradient Boosting Machines

Questions 15

Which of the following statements describes a Spark ML estimator?

Options:

An estimator is a hyperparameter arid that can be used to train a model

An estimator chains multiple alqorithms toqether to specify an ML workflow

An estimator is a trained ML model which turns a DataFrame with features into a DataFrame with predictions

An estimator is an alqorithm which can be fit on a DataFrame to produce a Transformer

An estimator is an evaluation tool to assess to the quality of a model

Buy Now

Questions 16

A data scientist has written a feature engineering notebook that utilizes the pandas library. As the size of the data processed by the notebook increases, the notebook's runtime is drastically increasing, but it is processing slowly as the size of the data included in the process increases.

Which of the following tools can the data scientist use to spend the least amount of time refactoring their notebook to scale with big data?

Options:

PySpark DataFrame API

pandas API on Spark

Spark SQL

Feature Store

Buy Now

Questions 17

A machine learning engineer is trying to perform batch model inference. They want to get predictions using the linear regression model saved at the pathmodel_urifor the DataFramebatch_df.

batch_dfhas the following schema:

customer_id STRING

The machine learning engineer runs the following code block to perform inference onbatch_dfusing the linear regression model atmodel_uri:

Databricks-Machine-Learning-Associate Question 17

In which situation will the machine learning engineer’s code block perform the desired inference?

Options:

When the Feature Store feature set was logged with the model at model_uri

When all of the features used by the model at model_uri are in a Spark DataFrame in the PySpark

When the model at model_uri only uses customer_id as a feature

This code block will not perform the desired inference in any situation.

When all of the features used by the model at model_uri are in a single Feature Store table

Buy Now

Questions 18

A data scientist is developing a single-node machine learning model. They have a large number of model configurations to test as a part of their experiment. As a result, the model tuning process takes too long to complete. Which of the following approaches can be used to speed up the model tuning process?

Options:

Implement MLflow Experiment Tracking

Scale up with Spark ML

Enable autoscaling clusters

Parallelize with Hyperopt

Buy Now

Questions 19

A data scientist is wanting to explore the Spark DataFrame spark_df. The data scientist wants visual histograms displaying the distribution of numeric features to be included in the exploration.

Which of the following lines of code can the data scientist run to accomplish the task?

Options:

spark_df.describe()

dbutils.data(spark_df).summarize()

This task cannot be accomplished in a single line of code.

spark_df.summary()

dbutils.data.summarize (spark_df)

Buy Now

Questions 20

A data scientist has developed a random forest regressor rfr and included it as the final stage in a Spark MLPipeline pipeline. They then set up a cross-validation process with pipeline as the estimator in the following code block:

Databricks-Machine-Learning-Associate Question 20

Which of the following is a negative consequence of includingpipelineas the estimator in the cross-validation process rather thanrfras the estimator?

Options:

The process will have a longer runtime because all stages of pipeline need to be refit or retransformed with each mode

The process will leak data from the training set to the test set during the evaluation phase

The process will be unable to parallelize tuning due to the distributed nature of pipeline

The process will leak data prep information from the validation sets to the training sets for each model

Buy Now

Questions 21

A data scientist is using the following code block to tune hyperparameters for a machine learning model:

Databricks-Machine-Learning-Associate Question 21

Which change can they make the above code block to improve the likelihood of a more accurate model?

Options:

Increase num_evals to 100

Change fmin() to fmax()

Change sparkTrials() to Trials()

Change tpe.suggest to random.suggest

Buy Now

Questions 22

A new data scientist has started working on an existing machine learning project. The project is a scheduled Job that retrains every day. The project currently exists in a Repo in Databricks. The data scientist has been tasked with improving the feature engineering of the pipeline’s preprocessing stage. The data scientist wants to make necessary updates to the code that can be easily adopted into the project without changing what is being run each day.

Which approach should the data scientist take to complete this task?

Options:

They can create a new branch in Databricks, commit their changes, and push those changes to the Git provider.

They can clone the notebooks in the repository into a Databricks Workspace folder and make the necessary changes.

They can create a new Git repository, import it into Databricks, and copy and paste the existing code from the original repository before making changes.

They can clone the notebooks in the repository into a new Databricks Repo and make the necessary changes.

Buy Now

Exam Code: Databricks-Machine-Learning-Associate

Exam Name: Databricks Certified Machine Learning Associate Exam

Last Update: Jun 21, 2025

Questions: 74

PDF + Testing Engine

$66 ~~$164.99~~

Testing Engine

$50 ~~$124.99~~

PDF (Q&A)

$42 ~~$104.99~~

buy now Databricks-Machine-Learning-Associate pdf

Databricks-Machine-Learning-Associate Databricks Certified Machine Learning Associate Exam Questions and Answers

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

Options:

Answer:

Explanation:

PDF + Testing Engine

Testing Engine

PDF (Q&A)

Quick Links

Why Us

Unlimited Packages

Marks4sure

Site Secure