ARTiBA Artificial Intelligence Engineer (AiE®) — Questions and Answers
Question 1: Which loss function is typically used for multi-class classification problems?
- Categorical Cross-Entropy (Correct answer)
- Binary Cross-Entropy
- Huber Loss
- Mean Squared Error
Correct answer: Categorical Cross-Entropy
Categorical cross-entropy measures the divergence between predicted class probabilities and the one-hot true labels across multiple classes.
Question 2: What term describes the algorithms that learn and make predictions from data without being explicitly programmed?
- Rule-based algorithms
- Pre-trained models
- Machine learning algorithms (Correct answer)
- Supervised learning
Correct answer: Machine learning algorithms
Machine learning algorithms are a subset of AI that enable systems to automatically learn and improve from experience without being explicitly programmed for every task. These algorithms identify patterns in data and use them to make predictions or decisions. They are distinct from rule-based systems, which rely on predefined explicit instructions.
Question 3: What distinguishes a 'high-risk AI system' under the EU AI Act?
- AI systems developed outside the EU
- AI used in domains like hiring, credit, healthcare, or law enforcement where errors have significant societal impact (Correct answer)
- AI systems with more than 1 billion parameters
- AI systems that run on public cloud infrastructure
Correct answer: AI used in domains like hiring, credit, healthcare, or law enforcement where errors have significant societal impact
The EU AI Act classifies AI as high-risk based on use-case domain (e.g., employment, critical infrastructure, law enforcement) due to potential for significant harm.
Question 4: What is the primary advantage of ensemble methods like Random Forest over a single decision tree?
- They are easier to interpret and visualize
- They require less computational power to train
- They eliminate the need for feature engineering
- They reduce variance by aggregating predictions from multiple models (Correct answer)
Correct answer: They reduce variance by aggregating predictions from multiple models
Random Forest reduces variance by averaging predictions from many decision trees trained on random subsets of data and features, resulting in better generalization than a single tree which tends to overfit.
Question 5: Which feature of Azure AI Language allows you to build a model that classifies text into your own custom categories?
- Conversational Language Understanding
- Named Entity Recognition
- Custom Text Classification (Correct answer)
- Text Analytics for Health
Correct answer: Custom Text Classification
Custom Text Classification lets you train a model on labeled examples to categorize documents into categories you define.
Question 6: What does the ROC-AUC score measure in a binary classification model?
- The average squared difference between predicted probabilities and true labels
- The probability that the model ranks a randomly chosen positive example higher than a randomly chosen negative example (Correct answer)
- The proportion of predictions that match the true class label across all thresholds
- The ratio of true positives to the total number of actual positive examples
Correct answer: The probability that the model ranks a randomly chosen positive example higher than a randomly chosen negative example
AUC (Area Under the ROC Curve) represents the probability that the model assigns a higher predicted probability to a random positive instance than to a random negative instance — an AUC of 1.0 is perfect, 0.5 is random.
Question 7: What is 'zero-shot prompting' when using an LLM?
- Using a model that has never been fine-tuned
- Training the model on zero labeled examples
- Prompting the model with zero tokens
- Asking the model to perform a task without providing any examples in the prompt (Correct answer)
Correct answer: Asking the model to perform a task without providing any examples in the prompt
Zero-shot prompting asks the LLM to complete a task using only instructions, relying entirely on knowledge from pretraining without in-context examples.
Question 8: Which of the following is an example of a generative machine learning model?
- Support Vector Machine
- Variational Autoencoder (VAE) (Correct answer)
- Logistic Regression
- Random Forest
Correct answer: Variational Autoencoder (VAE)
A Variational Autoencoder is a generative model that learns a latent distribution of the training data and can sample from it to generate new data points, unlike discriminative models that only learn class boundaries.
Question 9: What is 'prompt engineering' when working with LLMs?
- Optimizing LLM inference speed
- Training the LLM on a new dataset
- Designing input text to guide LLM behavior without modifying model weights (Correct answer)
- Writing code to deploy LLMs
Correct answer: Designing input text to guide LLM behavior without modifying model weights
Prompt engineering involves crafting input instructions, examples, and context to elicit desired outputs from an LLM without changing its parameters.
Question 10: What is the purpose of 'semantic ranking' in Azure Cognitive Search?
- It enforces security trimming based on user roles
- It replaces full-text keyword search with vector-only similarity
- It applies geospatial scoring to boost nearby results
- It uses language model understanding to re-rank results by semantic relevance beyond keyword matching (Correct answer)
Correct answer: It uses language model understanding to re-rank results by semantic relevance beyond keyword matching
Semantic ranking applies a transformer-based model to re-rank the top keyword search results by their true semantic relevance to the query.
Question 11: Which graph database query language is most commonly used to query property graphs in knowledge systems?
- Cypher (Correct answer)
- SOQL
- GraphQL
- SPARQL
Correct answer: Cypher
Cypher is the declarative query language used by Neo4j and other property graph databases, designed specifically for pattern matching over graph structures.
Question 12: In gradient descent, what does the learning rate control?
- The proportion of data used in each training batch
- The number of training epochs before the model converges
- The threshold for classifying a prediction as positive
- The size of the steps taken toward the minimum of the loss function (Correct answer)
Correct answer: The size of the steps taken toward the minimum of the loss function
The learning rate determines how large each update step is during gradient descent — too large causes oscillation or divergence, while too small results in slow convergence or getting stuck in local minima.
Question 13: Which principle of responsible AI ensures that stakeholders can understand and interrogate how an AI system makes decisions?
- Data compression
- Scalability
- Throughput optimization
- Explainability and transparency (Correct answer)
Correct answer: Explainability and transparency
Explainability means AI decisions can be understood by humans, while transparency ensures the system's design, data, and limitations are disclosed.
Question 14: What is 'AI red-teaming'?
- A technique for reducing LLM inference cost
- Adversarial testing where experts try to find failures, vulnerabilities, and harmful behaviors in AI systems (Correct answer)
- A competitive AI engineering tournament
- Training AI models on red-flagged data
Correct answer: Adversarial testing where experts try to find failures, vulnerabilities, and harmful behaviors in AI systems
AI red-teaming involves deliberately probing a system for safety failures, harmful outputs, or exploitable behaviors before deployment, mimicking adversarial users.
Question 15: Which of the following statements best describes a naive Bayes classifier?
- It partitions the feature space into rectangular regions using a series of threshold decisions
- It builds a decision boundary by maximizing the margin between classes
- It iteratively adjusts weights to minimize the cross-entropy loss between predictions and labels
- It applies Bayes' theorem assuming conditional independence between features given the class label (Correct answer)
Correct answer: It applies Bayes' theorem assuming conditional independence between features given the class label
Naive Bayes applies Bayes' theorem to compute the posterior probability of each class and classifies based on the highest probability, making the 'naive' assumption that features are conditionally independent given the class.
Question 16: In the context of decision trees, what is information gain?
- The total number of correctly classified examples after a split
- The reduction in entropy (uncertainty) in the target variable achieved by splitting on a given feature (Correct answer)
- The difference between the training accuracy and the validation accuracy at each tree depth
- The increase in model accuracy after adding a new feature to the training data
Correct answer: The reduction in entropy (uncertainty) in the target variable achieved by splitting on a given feature
Information gain measures how much a feature split reduces entropy (disorder) in the target variable — features with the highest information gain are chosen as split points to build a tree that separates classes most effectively.
Question 17: Which metric measures how well an LLM predicts a test corpus, with lower values indicating better language modeling?
- Precision
- AUC-ROC
- BLEU
- Perplexity (Correct answer)
Correct answer: Perplexity
Perplexity measures the exponentiated average negative log-likelihood of a test set; lower perplexity means the model assigns higher probability to the observed text.
Question 18: One of these states defines a problem in a search space.
- All of the mentioned
- Intermediate state
- Last state
- Initial state (Correct answer)
Correct answer: Initial state
In the context of a search space, the initial state is the fundamental starting point or configuration from which the search algorithm begins its exploration. It defines the problem's beginning before any actions are taken to reach a goal state. Without a defined initial state, the search cannot commence.
Question 19: What is 'embedding' in NLP, as used by transformer models?
- The process of tokenizing input text
- A dense, fixed-size vector representation of a token or text that captures semantic meaning (Correct answer)
- A technique for compressing the attention matrix
- Storing model weights in a database
Correct answer: A dense, fixed-size vector representation of a token or text that captures semantic meaning
Embeddings map discrete tokens (or entire texts) to dense vectors in a continuous space where semantically similar items are geometrically close.
Question 20: Which Azure service provides pre-built AI models for document understanding, including invoice and receipt extraction, without custom training?
- Azure Bot Service
- Azure Machine Learning AutoML
- Azure Form Recognizer (Document Intelligence) (Correct answer)
- Azure Custom Vision
Correct answer: Azure Form Recognizer (Document Intelligence)
Azure AI Document Intelligence includes pre-built models for common document types like invoices, receipts, and ID documents that work out of the box.
Question 21: What is the term for the process where a machine learning model generalizes well to new, unseen data?
- Overfitting
- Generalization (Correct answer)
- Regression
- Underfitting
Correct answer: Generalization
Generalization refers to a machine learning model's ability to perform accurately on new, previously unseen data, rather than just the data it was trained on. A model that generalizes well has learned the underlying patterns of the data without memorizing specific examples. This is a crucial indicator of a model's real-world applicability and effectiveness.
Question 22: What is RLHF (Reinforcement Learning from Human Feedback) used for in LLM development?
- Aligning LLM outputs with human preferences and reducing harmful outputs (Correct answer)
- Expanding the model's vocabulary
- Compressing the model for edge deployment
- Speeding up pretraining
Correct answer: Aligning LLM outputs with human preferences and reducing harmful outputs
RLHF uses human preference ratings to train a reward model, then fine-tunes the LLM with RL to produce outputs humans rate as better and safer.
Question 23: What is a vector database's primary role in an LLM application architecture?
- Storing model weights for fast loading
- Storing and querying high-dimensional embeddings for semantic similarity search (Correct answer)
- Caching LLM API responses
- Managing prompt templates
Correct answer: Storing and querying high-dimensional embeddings for semantic similarity search
Vector databases (e.g., Pinecone, Weaviate, Chroma) store embeddings and support fast approximate nearest-neighbor search for retrieval in RAG pipelines.
Question 24: What is 'fine-tuning' a large language model (LLM)?
- Prompting an LLM with few-shot examples
- Compressing an LLM using quantization
- Training an LLM from scratch on domain data
- Continuing training of a pretrained LLM on task-specific data to adapt its behavior (Correct answer)
Correct answer: Continuing training of a pretrained LLM on task-specific data to adapt its behavior
Fine-tuning updates the pretrained model's weights on a smaller, task-specific dataset, adapting general language understanding to a specific domain or task.
Question 25: What is 'adversarial robustness' in AI system design?
- Testing models under high concurrency
- Using multiple competing models in an ensemble
- A model's ability to maintain correct behavior when inputs are deliberately manipulated to cause errors (Correct answer)
- Training models on adversarial datasets
Correct answer: A model's ability to maintain correct behavior when inputs are deliberately manipulated to cause errors
Adversarial robustness measures how well a model resists adversarial examples — subtly perturbed inputs crafted to fool the model while appearing normal to humans.
Question 26: What is 'fairness through unawareness' and why is it insufficient?
- Using equal sample sizes per group; insufficient because it ignores distribution differences
- Excluding protected attributes (e.g., race) from model inputs; insufficient because proxies still encode the information (Correct answer)
- Balancing classes in training data; insufficient because it doesn't address test-time bias
- Ignoring all training data to avoid bias; insufficient because it produces random outputs
Correct answer: Excluding protected attributes (e.g., race) from model inputs; insufficient because proxies still encode the information
Excluding protected attributes doesn't prevent discrimination because correlated proxy variables (zip code, name) still allow the model to infer and act on them.
Question 27: What is the kernel trick in Support Vector Machines (SVM)?
- A pruning strategy that removes irrelevant features before training
- A method to reduce training time by approximating the decision boundary
- A regularization method that penalizes support vectors far from the margin
- A technique that implicitly maps data to a higher-dimensional space to find a linear separator (Correct answer)
Correct answer: A technique that implicitly maps data to a higher-dimensional space to find a linear separator
The kernel trick computes dot products in a higher-dimensional feature space without explicitly transforming the data, enabling SVMs to find linear decision boundaries for data that is nonlinearly separable in the original space.
Question 28: What is the purpose of a validation set during model training?
- To provide extra training samples
- To store the final model weights
- To replace the need for a test set
- To tune hyperparameters without touching the test set (Correct answer)
Correct answer: To tune hyperparameters without touching the test set
The validation set allows engineers to tune hyperparameters and detect overfitting before final evaluation on the test set.
Question 29: What does 'consent' mean in the context of using personal data to train AI models?
- Signing an NDA with data providers
- Obtaining informed agreement from individuals before using their personal data for model training (Correct answer)
- Having executives approve the training dataset
- Getting approval from the AI ethics board
Correct answer: Obtaining informed agreement from individuals before using their personal data for model training
Consent requires that individuals knowingly agree to how their personal data will be used, including for AI training, per regulations like GDPR and CCPA.
Question 30: Which evaluation metric is most appropriate when classes are severely imbalanced and the cost of false negatives is high?
- F1 Score (Correct answer)
- Accuracy
- Mean Squared Error
- R-squared
Correct answer: F1 Score
The F1 Score is the harmonic mean of precision and recall, making it suitable for imbalanced datasets where accuracy is misleading — it balances the cost of false positives and false negatives.
Question 31: In k-means clustering, how is the optimal number of clusters (k) typically determined?
- By running the algorithm with k=1 and incrementally adding clusters until accuracy stops improving
- By using the default value of k=3 as it works for most datasets
- By using the elbow method, which plots inertia vs. k and identifies the point of diminishing returns (Correct answer)
- By setting k equal to the square root of the number of data points
Correct answer: By using the elbow method, which plots inertia vs. k and identifies the point of diminishing returns
The elbow method plots the within-cluster sum of squares (inertia) against different values of k — the 'elbow' point where adding more clusters yields diminishing reductions in inertia suggests the optimal k.
Question 32: Why is feature standardization (zero mean, unit variance) important before applying algorithms like SVM or k-nearest neighbors?
- It reduces the number of training iterations required for convergence
- It ensures the model produces probabilistic outputs between 0 and 1
- It prevents features with larger numerical ranges from dominating distance-based calculations (Correct answer)
- It automatically removes irrelevant features from the dataset
Correct answer: It prevents features with larger numerical ranges from dominating distance-based calculations
Distance-based algorithms are sensitive to feature scale — a feature ranging 0–10,000 will dominate a feature ranging 0–1 in distance calculations, so standardization puts all features on equal footing.
Question 33: What does 'temperature' control in LLM text generation?
- The model's confidence threshold for answering
- The maximum length of generated text
- The randomness of token sampling — higher values produce more diverse outputs (Correct answer)
- The computational load during inference
Correct answer: The randomness of token sampling — higher values produce more diverse outputs
Temperature scales the logits before softmax; higher values flatten the distribution (more random), lower values sharpen it (more deterministic).
Question 34: Which of the following best describes an embedding in AI?
- A compressed image file format
- A one-hot encoded categorical variable
- A dense vector representation of data in a continuous space (Correct answer)
- A rule-based lookup table
Correct answer: A dense vector representation of data in a continuous space
Embeddings map discrete objects (words, items) to dense vectors so that similar objects are close together in vector space.
Question 35: What is the primary difference between a parametric and a non-parametric machine learning model?
- Parametric models use gradient descent, while non-parametric models use Bayesian inference
- Parametric models require GPU acceleration, while non-parametric models run on CPUs
- Parametric models can only handle continuous features, while non-parametric models handle categorical features
- Parametric models assume a fixed functional form with a set number of parameters, while non-parametric models grow in complexity with the training data (Correct answer)
Correct answer: Parametric models assume a fixed functional form with a set number of parameters, while non-parametric models grow in complexity with the training data
Parametric models (e.g., linear regression, logistic regression) summarize data with a fixed number of parameters regardless of dataset size, while non-parametric models (e.g., k-NN, kernel SVM) retain training data and grow in complexity with more data.
Question 36: In AI system design, what is the purpose of a 'human-in-the-loop' mechanism?
- To replace AI with human workers
- To include human oversight and intervention in AI decision pipelines for critical or uncertain cases (Correct answer)
- To monitor server infrastructure
- To train models using human-labeled data
Correct answer: To include human oversight and intervention in AI decision pipelines for critical or uncertain cases
Human-in-the-loop keeps humans involved in reviewing, correcting, or approving AI decisions, especially where errors have high stakes.
Question 37: What is the process of adjusting a machine learning model's parameters to minimize errors on the training data?
- Optimization (Correct answer)
- Overfitting
- Regression
- Underfitting
Correct answer: Optimization
Optimization is the process of iteratively adjusting a machine learning model's internal parameters (like weights and biases) to minimize a defined error or loss function. This process aims to find the best set of parameters that allows the model to make the most accurate predictions on the training data. Algorithms like gradient descent are commonly used for this purpose.
Question 38: What does 'tokenization' mean in the context of NLP preprocessing?
- Removing stopwords from a document
- Encrypting text data before storage
- Splitting raw text into discrete units (tokens) for model input (Correct answer)
- Converting text to audio
Correct answer: Splitting raw text into discrete units (tokens) for model input
Tokenization breaks raw text into tokens (words, subwords, or characters) that are then mapped to numeric IDs for model processing.
Question 39: What is the primary challenge addressed by 'federated learning' in AI?
- Distributing training compute across multiple GPUs
- Federating API access to AI models
- Training models across multiple cloud providers
- Training AI models across decentralized devices without sharing raw data, preserving privacy (Correct answer)
Correct answer: Training AI models across decentralized devices without sharing raw data, preserving privacy
Federated learning trains models locally on each device and aggregates only model updates (not raw data), enabling collaboration without centralizing sensitive data.
Question 40: Which pricing tier of Azure Cognitive Services provides a Service Level Agreement (SLA) guarantee?
- Standard (S) tier (Correct answer)
- Neither tier includes an SLA
- Both Free and Standard tiers
- Free (F0) tier
Correct answer: Standard (S) tier
Only the Standard (S) paid tier of Azure Cognitive Services includes SLA commitments; the Free tier does not.
Question 41: What does 'scalable oversight' aim to solve in AI alignment research?
- Automating model retraining at scale
- Scaling model parameters beyond 1 trillion
- Maintaining meaningful human supervision of AI as systems become more capable than the humans overseeing them (Correct answer)
- Distributing AI workloads across clusters
Correct answer: Maintaining meaningful human supervision of AI as systems become more capable than the humans overseeing them
Scalable oversight addresses how to keep humans in meaningful control of AI decisions when the AI may eventually be more capable than the humans evaluating it.
Question 42: Which parameter-efficient fine-tuning technique adds low-rank decomposition matrices to model layers instead of updating all weights?
- RLHF (Reinforcement Learning from Human Feedback)
- LoRA (Low-Rank Adaptation) (Correct answer)
- Knowledge distillation
- Full fine-tuning
Correct answer: LoRA (Low-Rank Adaptation)
LoRA freezes pretrained weights and injects trainable low-rank matrices, drastically reducing the number of parameters updated during fine-tuning.
Question 43: Which regularization technique adds the sum of the absolute values of model weights as a penalty term to the loss function?
- L1 (Lasso) regularization (Correct answer)
- Dropout regularization
- Elastic Net regularization
- L2 (Ridge) regularization
Correct answer: L1 (Lasso) regularization
L1 (Lasso) regularization penalizes the sum of absolute values of weights, which encourages sparsity by driving some weights to exactly zero, effectively performing feature selection.
Question 44: What is 'model governance' in enterprise AI system design?
- Optimizing model training compute costs
- Version control for model code
- Policies, controls, and oversight processes that manage the AI model lifecycle from development to retirement (Correct answer)
- A software library for model deployment
Correct answer: Policies, controls, and oversight processes that manage the AI model lifecycle from development to retirement
Model governance encompasses approval workflows, audit trails, risk assessments, and accountability structures ensuring AI models are developed and used responsibly.
Question 45: What is 'early stopping' as a regularization technique in deep learning?
- Freezing certain layers during training
- Reducing learning rate after a fixed number of epochs
- Stopping training when loss reaches zero
- Halting training when validation loss stops improving to prevent overfitting (Correct answer)
Correct answer: Halting training when validation loss stops improving to prevent overfitting
Early stopping monitors validation loss and stops training when it starts increasing, saving the model at the point of best generalization.
Question 46: What is the 'right to explanation' under AI regulations like the EU AI Act?
- The requirement to open-source AI training code
- The obligation to explain AI research publicly
- The right for individuals to receive a meaningful explanation of automated decisions that affect them (Correct answer)
- The right for AI engineers to document their models
Correct answer: The right for individuals to receive a meaningful explanation of automated decisions that affect them
Regulations like the EU AI Act and GDPR grant individuals the right to understand why an automated system made a decision affecting them.
Question 47: What does the bias-variance tradeoff describe in machine learning?
- The ratio of labeled to unlabeled data in a dataset
- The relationship between training speed and model accuracy
- The balance between model complexity and generalization to unseen data (Correct answer)
- The tradeoff between precision and recall in classification
Correct answer: The balance between model complexity and generalization to unseen data
The bias-variance tradeoff describes how increasing model complexity reduces bias but increases variance, while simpler models have high bias but low variance — the goal is to minimize total error on unseen data.
Question 48: What is the context window in a large language model?
- The UI window showing model outputs
- The time window used for model training
- The layer of attention that focuses on the current word
- The maximum number of tokens the model can process in a single input/output sequence (Correct answer)
Correct answer: The maximum number of tokens the model can process in a single input/output sequence
The context window defines the maximum number of tokens (input + output) an LLM can process at once, limiting how much text it can consider.
Question 49: What is the role of 'reranking' in a two-stage retrieval pipeline?
- Re-scoring a small candidate set using a higher-capacity model to improve final ranking quality (Correct answer)
- Replacing the first-stage retriever entirely with a more powerful model
- Filtering documents that exceed the context window length
- Randomly shuffling retrieved documents to reduce bias
Correct answer: Re-scoring a small candidate set using a higher-capacity model to improve final ranking quality
A reranker (e.g., a cross-encoder) takes the top-k candidates from a fast first-stage retriever and applies a more computationally expensive relevance model to produce a higher-quality final ranking.
Question 50: What is the primary difference between gradient boosting and bagging (e.g., Random Forest)?
- Bagging is only applicable to regression tasks, while gradient boosting handles classification
- Bagging builds models sequentially, while gradient boosting builds them in parallel
- Gradient boosting averages predictions, while bagging uses voting to select the best model
- Gradient boosting builds models sequentially where each model corrects the errors of the previous one (Correct answer)
Correct answer: Gradient boosting builds models sequentially where each model corrects the errors of the previous one
Gradient boosting builds an ensemble sequentially, where each new model focuses on correcting the residual errors of the combined previous models, reducing bias — unlike bagging which builds models independently in parallel to reduce variance.
Question 51: What is an 'AI incident' as defined in responsible AI frameworks?
- A drop in model accuracy below a threshold
- An event where an AI system causes or contributes to harm, near-miss, or unexpected negative consequences in deployment (Correct answer)
- A disagreement between AI engineers about model architecture
- A model failing to converge during training
Correct answer: An event where an AI system causes or contributes to harm, near-miss, or unexpected negative consequences in deployment
An AI incident is any real-world event where an AI system causes harm or poses significant risk, tracked in repositories like the AI Incident Database.
Question 52: What is the primary purpose of k-fold cross-validation?
- To obtain a more reliable estimate of model performance by using all data for both training and validation (Correct answer)
- To increase the size of the training dataset through data augmentation
- To reduce training time by splitting data into smaller batches
- To select the optimal number of features for a model
Correct answer: To obtain a more reliable estimate of model performance by using all data for both training and validation
K-fold cross-validation splits data into k subsets, trains on k-1 folds, and validates on the remaining fold, rotating until all folds are used — giving a robust performance estimate without wasting data.
Question 53: What is 'hallucination' in the context of LLMs?
- The model outputting garbled text due to tokenization errors
- The model producing plausible-sounding but factually incorrect or fabricated information (Correct answer)
- The model refusing to answer sensitive questions
- The model generating extremely long outputs
Correct answer: The model producing plausible-sounding but factually incorrect or fabricated information
LLM hallucination refers to the model confidently generating false, invented information not grounded in training data or retrieved context.
Question 54: What does 'knowledge grounding' refer to in natural language processing?
- Converting knowledge graphs to relational databases
- Tokenizing text into subword units
- Linking language model outputs to verified external facts or real-world entities (Correct answer)
- Initializing model weights with domain-specific values
Correct answer: Linking language model outputs to verified external facts or real-world entities
Knowledge grounding connects model-generated statements to external, verifiable sources to reduce hallucinations and improve factual reliability.
Question 55: What distinguishes a generative adversarial network (GAN) from a variational autoencoder (VAE)?
- GANs use an adversarial training loop between generator and discriminator; VAEs optimize a variational lower bound (Correct answer)
- GANs encode data to latent space; VAEs generate from noise
- GANs require labeled data; VAEs do not
- GANs use supervised learning; VAEs use unsupervised learning
Correct answer: GANs use an adversarial training loop between generator and discriminator; VAEs optimize a variational lower bound
GANs pit a generator against a discriminator in an adversarial game, while VAEs learn a probabilistic latent space by maximizing an evidence lower bound (ELBO).
Question 56: Your Azure Bot sends proactive messages to users in Microsoft Teams. Which Bot Framework feature must be stored and reused to send proactive messages?
- The conversation reference obtained from a previous turn (Correct answer)
- The user's Azure AD object ID
- The bot's App ID and password
- The Teams channel webhook URL
Correct answer: The conversation reference obtained from a previous turn
A conversation reference captured during an active turn contains the service URL and conversation ID needed to initiate proactive messages later.
Question 57: What is 'chain-of-thought prompting' in LLMs?
- Chaining multiple LLM API calls sequentially
- Using multiple models in a pipeline
- Training the model on logical reasoning datasets
- Prompting the model to reason through intermediate steps before giving a final answer (Correct answer)
Correct answer: Prompting the model to reason through intermediate steps before giving a final answer
Chain-of-thought prompting encourages LLMs to explicitly generate reasoning steps, significantly improving performance on complex reasoning tasks.
Question 58: What is 'semantic chunking' in the context of building RAG systems?
- Encrypting document chunks before embedding
- Splitting documents by fixed token counts
- Filtering out low-quality document segments
- Dividing documents into chunks based on semantic coherence to preserve meaningful context (Correct answer)
Correct answer: Dividing documents into chunks based on semantic coherence to preserve meaningful context
Semantic chunking splits documents at natural semantic boundaries rather than fixed sizes, improving retrieval quality by keeping coherent content together.
Question 59: Which deep learning architecture introduced skip (residual) connections to enable training of very deep networks?
- VGG
- AlexNet
- ResNet (Correct answer)
- GoogLeNet
Correct answer: ResNet
ResNet introduced residual connections that add layer inputs directly to outputs, allowing gradients to flow more easily and enabling hundreds of layers.
Question 60: What is the 'attention mechanism' in transformer-based models?
- A mechanism that computes weighted relationships between all positions in a sequence (Correct answer)
- A technique for data augmentation in NLP
- A regularization method for language models
- A method to prune unimportant neurons
Correct answer: A mechanism that computes weighted relationships between all positions in a sequence
Attention computes a weighted sum of values based on query-key similarity, allowing the model to focus on relevant parts of the input regardless of distance.
ARTiBA Artificial Intelligence Engineer (AiE®)
The AiE® certification validates expertise across core AI engineering domains including machine learning, neural networks, NLP, and responsible AI deployment, aligned with the AMDEX™ Knowledge Framework. It is a globally recognized credential for professionals building and implementing AI systems.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds