Explainable LLMs That Cite Source Evidence
1. Core Principles of Explainability in AI
Core Principles of Explainability in AI
Explainability in AI refers to the ability of a model to provide human-understandable justifications for its decisions. For large language models (LLMs), this involves not only generating coherent outputs but also tracing the reasoning process back to verifiable sources. The core principles of explainability can be decomposed into three foundational pillars: transparency, interpretability, and attribution.
Transparency
Transparency requires that the internal mechanisms of an AI system be accessible for inspection. For LLMs, this means exposing the model's architecture, training data distribution, and decision pathways. A transparent model allows researchers to audit:
- Weight matrices and attention patterns
- Training data provenance and potential biases
- Fine-tuning procedures and hyperparameter choices
Mathematically, transparency can be quantified through measures like parameter saliency, which identifies influential weights in the network. For a given input x and output y, the saliency S of parameter θi is:
Interpretability
Interpretability focuses on making model outputs understandable to human users. For LLMs that cite sources, this involves:
- Generating natural language explanations alongside predictions
- Highlighting relevant passages from source documents
- Providing confidence estimates for factual claims
A key technique is attention visualization, which shows how input tokens influence output generation. The attention weight αij between token i and j can be represented as:
where Q, K are query and key matrices, and dk is the dimension of the key vectors.
Attribution
Attribution links model outputs to specific evidence in the training corpus. Advanced methods include:
- Neural memory networks that store and retrieve source snippets
- Gradient-based feature importance scoring
- Counterfactual analysis showing how outputs change with modified inputs
The attribution strength A of a source document D for prediction y can be computed via integrated gradients:
where x represents the model's internal representation of document D, and x' is a baseline representation.
Practical Implementation
Modern explainable LLMs implement these principles through hybrid architectures combining:
- Retrieval-augmented generation (RAG) for source grounding
- Attention rollout techniques for decision tracing
- Contrastive explanations that compare against alternative outputs
The effectiveness of these methods is evaluated using metrics like:
where Ei is the explanation, Mi is the model output, and human() measures understandability on a 0-1 scale.

Challenges in Interpreting LLM Outputs
Large Language Models (LLMs) generate text by predicting the next token in a sequence based on learned statistical patterns. While this enables fluent and coherent outputs, it introduces several challenges for interpretability and verification of source evidence. The probabilistic nature of LLMs means their responses are not deterministic, making it difficult to trace the reasoning behind specific outputs.
Lack of Explicit Reasoning Traces
Unlike rule-based systems, LLMs do not maintain explicit symbolic representations of their reasoning process. The internal computations occur in high-dimensional vector spaces through attention mechanisms and feedforward networks, making it challenging to extract human-understandable justifications. For example, when an LLM answers a factual question, the model does not explicitly retrieve or cite a specific source—it generates text based on patterns observed during training.
Here, the probability distribution over the next token \( w_t \) is computed from the hidden state \( \mathbf{h}_{t-1} \), which encodes contextual information but lacks interpretable structure.
Hallucinations and Confidence Miscalibration
LLMs frequently generate plausible but incorrect or unsupported statements (hallucinations) with high confidence. This occurs because the training objective maximizes likelihood over a corpus rather than factual accuracy. The model's confidence scores—often derived from softmax probabilities—do not reliably indicate correctness, as they reflect frequency-based priors rather than epistemic uncertainty.
Attention Weights as Poor Proxies for Explanation
While attention mechanisms highlight which input tokens influence specific outputs, these weights are noisy and often fail to correlate with human notions of relevance. Multi-head attention distributes information across numerous parallel layers, making it difficult to attribute outputs to specific input segments:
Empirical studies show that alternative attention patterns can yield similar outputs, suggesting that weights alone are insufficient for faithful explanations.
Training Data Memorization vs. Generalization
LLMs interpolate and recombine patterns from training data without explicitly distinguishing between memorized facts and generalized knowledge. This creates ambiguity when attempting to verify sources—a model may generate text matching a specific source without having direct access to it during inference. Differential privacy analyses reveal that even rare training examples can be reconstructed from model outputs.
Dependency on Prompt Formulation
LLM outputs are highly sensitive to prompt phrasing, temperature settings, and sampling strategies. Minor rewording can lead to contradictory responses, complicating reproducibility. For instance, a question framed as "What causes climate change?" may yield different citations than "List scientific consensus on climate change drivers."
Scalability of Verification
Real-time citation requires either:
- Retrieval-augmented architectures that query external databases, introducing latency-cost tradeoffs
- Post-hoc justification methods that search for supporting evidence after text generation
Both approaches struggle with combinatorial explosion when validating long-form outputs against potential sources.
Importance of Source Citation for Trustworthiness
Source citation in large language models (LLMs) is not merely a stylistic choice—it is a foundational requirement for establishing trustworthiness in AI-generated outputs. Without verifiable references, even the most accurate responses from an LLM remain suspect, as users lack the means to independently validate claims. This is particularly critical in domains like scientific research, legal analysis, and medical diagnostics, where incorrect or unsubstantiated information can have severe consequences.
Verifiability and Accountability
The primary value of source citation lies in enabling verifiability. When an LLM cites its sources, users can trace the origin of the information, assess the credibility of the reference, and confirm its accuracy. This process mirrors academic peer review, where claims must be backed by citable evidence. For example, if an LLM states that "the Higgs boson was discovered at CERN in 2012," citing the original ATLAS and CMS collaboration papers allows physicists to verify the claim against primary experimental data.
Mathematically, the trustworthiness T of an LLM response can be modeled as a function of source reliability S and transparency τ:
where α and β are weighting factors representing the relative importance of source quality versus transparency in the given context.
Mitigating Hallucinations
Source citation acts as a constraint mechanism against model hallucinations—fabricated information presented as fact. By requiring the model to ground its responses in existing references, the probability of generating unsupported claims decreases. This is especially relevant for open-ended queries where the model might otherwise extrapolate beyond its training data. For instance, a legal LLM citing specific case law (e.g., Roe v. Wade, 410 U.S. 113) provides a check against generating plausible but incorrect legal interpretations.
Domain-Specific Requirements
Different fields impose varying standards for source reliability:
- Medicine: Citations must point to peer-reviewed clinical studies or established medical guidelines (e.g., NIH, WHO publications).
- Law: References should include official case reporters, statutes, or regulatory documents with precise section numbering.
- Engineering: Technical specifications and standards (IEEE, ISO) carry more weight than secondary interpretations.
Failure to meet these domain-specific citation standards undermines the utility of LLM outputs for professional use. A medical diagnosis without references to UpToDate or PubMed sources, for example, would be ethically unacceptable for clinical decision support.
Audit Trails and Reproducibility
Source citations create an audit trail that enables reproducibility—a core principle of scientific inquiry. When an LLM's reasoning process is anchored to specific references, researchers can:
- Reconstruct the information pathway the model used to generate its response
- Identify potential biases in the source material
- Update outputs when new evidence supersedes old references
This is particularly valuable for longitudinal analyses where the state of knowledge evolves over time. A climate science model citing IPCC assessment reports from different years allows users to track how consensus positions have changed.
Legal and Ethical Compliance
Proper source attribution also addresses copyright and plagiarism concerns. When LLMs reproduce substantial portions of copyrighted material (e.g., journal articles, technical manuals), citations provide necessary attribution while falling under fair use provisions for educational purposes. This becomes legally significant when AI systems are used for commercial research or content generation.
2. Retrieval-Augmented Generation (RAG) Architectures
2.1 Retrieval-Augmented Generation (RAG) Architectures
Retrieval-Augmented Generation (RAG) combines dense retrieval with autoregressive language models to ground generations in external knowledge sources. The architecture consists of three key components: a retriever, an encoder, and a generator. Given an input query x, the retriever fetches relevant documents D = {d₁, d₂, ..., d_k} from a corpus, which are then encoded and passed to the generator alongside the original input.
Mathematical Formulation
The probability distribution over output tokens y_t at step t is conditioned on both the input x and retrieved documents D:
where P(d | x) represents the retriever's relevance score for document d, typically computed using maximum inner product search (MIPS) over dense embeddings:
Here, E_Q and E_D are query and document encoders respectively, often implemented as dual-encoder transformers.
Architecture Variants
Dense Retrieval
Modern RAG systems employ dense passage retrieval (DPR) using BERT-style encoders fine-tuned on question-answer pairs. The retriever is trained to maximize the likelihood of positive passages:
Fusion-in-Decoder
The generator processes retrieved documents through cross-attention mechanisms. The Fusion-in-Decoder approach concatenates all retrieved passages and attends to them jointly:
Practical Implementation
Production RAG systems face key engineering challenges:
- Latency-accuracy tradeoff: Larger retrieval corpora improve recall but increase search time
- Document chunking: Optimal passage segmentation balances context length with information density
- Asynchronous updates: Decoupling document index updates from model serving
Recent advances like ColBERTv2 introduce late interaction mechanisms that compute fine-grained relevance scores while maintaining efficient retrieval:
Evaluation Metrics
RAG systems require specialized evaluation beyond standard language modeling metrics:
- Citation recall: Percentage of generated claims supported by retrieved documents
- Answer faithfulness: Factual consistency between generations and sources
- Retrieval precision@k: Relevance of top-k retrieved passages

Attention Mechanisms for Evidence Localization
Transformer-based large language models (LLMs) rely on attention mechanisms to dynamically weight the relevance of input tokens when generating outputs. For explainable LLMs that cite sources, attention weights serve as a direct mechanism for evidence localization, revealing which parts of the input influenced a given prediction. The core mathematical formulation involves computing query-key-value attention scores:
Here, Q (queries), K (keys), and V (values) are learned linear transformations of the input embeddings, and dk is the dimension of the key vectors. The softmax operation normalizes the attention weights, allowing interpretable inspection of token contributions.
Multi-Head Attention for Fine-Grained Evidence
Multi-head attention extends this by parallelizing attention across h subspaces, each capturing distinct semantic relationships. For a model with h heads, the output is computed as:
where each head applies scaled dot-product attention independently. This allows the model to attend to different evidence spans simultaneously—e.g., one head might focus on factual entities while another tracks syntactic dependencies.
Cross-Attention for Retrieval-Augmented Models
In retrieval-augmented LLMs, cross-attention layers compute relevance between generated tokens and retrieved documents. Given a query q (current decoder state) and document tokens D, the evidence score for token di is:
where W is a learned projection matrix. These scores directly indicate which retrieved passages informed the model's output, enabling verifiable citations.
Practical Implementation Challenges
While attention weights provide a theoretically sound mechanism for evidence attribution, several practical issues arise:
- Attention head diversity: Not all heads learn interpretable patterns; some may capture noise or redundant features.
- Softmax saturation: Highly peaked distributions can obscure secondary evidence sources.
- Layer-wise propagation: Evidence must be traced through multiple transformer layers, as later layers may redistribute attention.
Recent approaches like attention rollout and gradient-based attribution complement raw attention weights to address these limitations. For example, gradient-weighted class activation mapping (Grad-CAM) for transformers computes:
where Ak are the attention maps and αkc are gradient-derived importance weights for class c.

2.3 Probabilistic Confidence Scoring of Citations
Large language models (LLMs) generate citations by retrieving relevant documents and attributing statements to them. However, not all citations are equally reliable—some may be tangential, weakly supported, or even incorrect. Probabilistic confidence scoring quantifies the strength of evidence behind each citation using statistical methods, enabling users to assess the trustworthiness of generated outputs.
Bayesian Formulation of Citation Confidence
The confidence score for a citation can be modeled as a posterior probability given the evidence. Let D be the retrieved document and S be the statement attributed to it. The confidence C is:
Applying Bayes' theorem, this decomposes into:
where the prior P(S is supported by D) can be estimated from document quality metrics, and the likelihood term evaluates how strongly the retrieval signals (e.g., semantic similarity, positional information) indicate support.
Evidence Aggregation from Multiple Signals
Modern systems combine multiple evidence signals through learned weighting:
where σ is the logistic function, wi are learned weights, and fi transforms raw signals like:
- Lexical overlap between statement and document spans
- Semantic similarity from embedding spaces
- Attention weights in cross-document transformer layers
- Positional information of cited text spans
Calibration and Uncertainty Quantification
Well-calibrated confidence scores should reflect true correctness probabilities. This is achieved through:
Minimizing this via temperature scaling or isotonic regression ensures that a citation with 0.8 confidence is correct 80% of the time. For uncertainty, we can compute:
which peaks at 0.5 confidence and decreases toward definitive (0 or 1) predictions.
Implementation Considerations
In practice, confidence scoring requires:
- Differentiable retrieval architectures that propagate gradient signals
- Human-annotated datasets of statement-document support judgments
- Regular evaluation on held-out factual verification tasks
State-of-the-art systems like Atlas and RARR achieve 85-90% AUC in distinguishing well-supported from poorly-supported citations using these methods.
3. Data Pipeline Design for Evidence Anchoring
Data Pipeline Design for Evidence Anchoring
Evidence Retrieval and Document Chunking
The first stage involves preprocessing source documents into retrievable chunks while preserving metadata. Documents are split using semantic boundaries (paragraphs, sections) rather than fixed token windows to maintain coherence. Each chunk is embedded using a contrastively trained encoder (e.g., ANCE or ColBERT) to enable dense retrieval. The chunking process must preserve:
- Document provenance (source URL, timestamp, author)
- Positional metadata (section hierarchy, adjacent chunks)
- Cross-references (citations, figure/table pointers)
Hierarchical Vector Indexing
For efficient retrieval across billion-scale corpora, we implement a two-tiered indexing strategy:
- Coarse-level: FAISS IVF indexes with product quantization for fast approximate search
- Fine-level: Exact nearest-neighbor search within candidate clusters using HNSW graphs
The indexing process optimizes for:
where \( R(q) \) represents ground truth relevant chunks and \( \text{NN}_k \) denotes retrieved neighbors.
Dynamic Reranking with Cross-Attention
Initial retrievals are refined using a lightweight cross-encoder that computes attention scores between query tokens and document chunks:
The final evidence score combines lexical overlap (BM25), semantic similarity (dense retrieval), and attention weights:
Evidence Attribution in Generation
During text generation, the LLM attends to both retrieved chunks and its internal knowledge. We modify the standard attention mechanism to track external evidence influence:
where \( \mathbf{e}_j \) represent retrieved evidence embeddings. Attribution scores are computed via gradient-based feature importance methods.
Versioned Evidence Tracking
To handle evolving knowledge, the pipeline implements:
- Timestamped document snapshots with incremental indexing
- Version-aware retrieval using temporal embeddings
- Change detection for automatic evidence updates
The complete pipeline achieves sub-200ms latency for end-to-end evidence retrieval and attribution while maintaining >90% precision on factual grounding tasks.

3.2 Training Protocols for Citation-Aware Models
Training large language models to generate citations requires specialized protocols that go beyond standard autoregressive pretraining. The key challenge lies in teaching the model to retrieve relevant sources, assess their reliability, and generate grounded textual output with proper attribution.
Multi-Task Learning Framework
Citation-aware models typically employ a multi-task objective combining:
- Language modeling loss (standard next-token prediction)
- Retrieval alignment loss (matching generated text to source passages)
- Citation prediction loss (identifying when/where to cite)
where α, β, γ are task weighting hyperparameters typically optimized via grid search.
Retrieval-Augmented Training
The training pipeline incorporates:
- Dense passage retrieval (DPR) during forward passes
- Hard negative mining to improve discrimination between relevant and irrelevant sources
- Dynamic memory banks that maintain recent retrieved contexts
The retrieval component is jointly trained with the language model using maximum inner product search (MIPS) optimization:
where Eq and Ed are query and document encoders respectively.
Citation Position Prediction
A critical subtask involves predicting when to insert citations. This is modeled as a binary classification head trained on:
- Lexical features (n-gram patterns preceding citations)
- Semantic features (BERT-style contextual embeddings)
- Structural features (position in paragraph/document)
The citation probability at token position t is computed as:
where ht is the hidden state at position t and k defines the local context window.
Source Reliability Estimation
Models are trained to assess source quality through:
- Domain-specific credibility signals (journal impact factors, citation counts)
- Internal consistency checks (agreement with other retrieved sources)
- Stance detection (identifying potential biases)
The reliability score R for source s is computed as:
where fmeta extracts metadata features and fcontent analyzes textual content.
Training Data Construction
High-quality training data requires:
- Expert-annotated citation spans in academic/scientific texts
- Diverse retrieval corpora covering multiple domains
- Adversarial examples containing misleading citations
The data pipeline typically involves:
def generate_citation_examples(text, sources):
# Step 1: Align source passages to text spans
alignments = find_textual_alignments(text, sources)
# Step 2: Generate positive and negative examples
positives = [(t, s) for t, s in alignments if is_valid_citation(t, s)]
negatives = [(t, s) for t, s in alignments if not is_valid_citation(t, s)]
# Step 3: Balance dataset
return balance_examples(positives, negatives)
Fine-Tuning Strategies
Final model optimization employs:
- Curriculum learning (gradually increasing citation density)
- Reinforcement learning (rewarding proper attribution)
- Adversarial training (against citation manipulation)
The RL reward function typically includes:
where the λ parameters control different aspects of citation quality.

3.3 Evaluation Metrics for Attribution Accuracy
Evaluating the attribution accuracy of explainable LLMs requires metrics that quantify how well generated citations align with ground-truth source evidence. Unlike traditional language model evaluation, attribution metrics must assess both the correctness of the generated text and the validity of its supporting references.
Precision and Recall for Source Attribution
The most fundamental metrics adapt information retrieval's precision and recall to measure citation quality. Given a set of ground-truth source spans S and predicted citations Ĉ:
These can be computed at different granularities - from exact token matches to fuzzy overlap using ROUGE or BERTScore. The F1 score combines both metrics:
Attribution-Aware Text Quality Metrics
Standard text generation metrics like BLEU or ROUGE fail to capture attribution correctness. Recent work proposes hybrid metrics:
- Attribution-BLEU: Modifies BLEU by downweighting n-grams without proper citations
- Cite-F1: Jointly scores text quality and citation accuracy using weighted harmonic mean
- Verification Probability: Uses an auxiliary model to estimate likelihood that cited sources support the claim
Hierarchical Evaluation for Multi-Span Attribution
When LLMs cite multiple evidence spans, we need hierarchical evaluation:
where N is the number of distinct claims in the generated text. This accounts for partial matches while preventing overcounting.
Human-Aligned Metrics
Automated metrics should correlate with human judgments. Common protocols include:
- Attribution Accuracy: Percentage of claims humans verify as properly cited
- Citation Necessity: Measures whether citations were needed for factual claims
- Hallucination Rate: Counts unsupported factual assertions despite citations
Recent benchmarks like AttributionQA and CiteBench provide standardized test sets with human annotations for these metrics.
Latency-Aware Evaluation
In production systems, attribution introduces computational overhead. Key operational metrics include:
where tretrieve measures source retrieval time and tverify measures verification time. The optimal tradeoff between accuracy and latency depends on application requirements.
4. Medical Diagnosis Systems with Literature References
Medical Diagnosis Systems with Literature References
Large language models (LLMs) deployed in medical diagnosis must provide not only accurate predictions but also verifiable evidence from trusted sources. The key challenge lies in aligning model outputs with peer-reviewed literature while maintaining clinical relevance. This requires three technical components: retrieval-augmented generation (RAG), source attribution mechanisms, and confidence calibration.
Retrieval-Augmented Generation Architecture
The RAG framework combines a dense retriever with a generative transformer. Given an input patient description x, the system first queries a medical corpus (e.g., PubMed, UpToDate) using:
where fθ and gϕ are dual encoders trained via contrastive learning on (query, document) pairs. The retrieved evidence r then conditions the LLM's generation:
Source Attribution Mechanisms
For traceability, the model must output citations in standard formats (e.g., AMA, Vancouver style). This is implemented through:
- Pointer networks that learn to select spans from retrieved documents
- BibTeX-aware decoding that constrains the output space to valid citation formats
- Attention-based attribution where citation locations are determined by cross-attention weights between generated text and source documents
Confidence Calibration
Medical applications require well-calibrated uncertainty estimates. We apply:
where α and β are learned parameters that scale and shift the model's logits. The temperature-scaling is trained on a held-out validation set of (input, evidence, expert judgment) triples.
Implementation Example
A deployed system might process the input "45yo male with crushing substernal chest pain radiating to left arm" and output:
The presentation is concerning for acute coronary syndrome (ACS) with likelihood 87% [1]. Immediate ECG and troponin testing are recommended.
The underlying architecture verifies that the cited paper actually contains supporting evidence for ACS diagnosis and that the confidence score aligns with the study's reported positive predictive value.
Evaluation Metrics
Performance is measured through:
- Citation precision/recall: Percentage of generated statements properly grounded in cited sources
- Clinical utility score: Expert-rated usefulness on a 1-5 Likert scale
- Hallucination rate: Instances where citations don't support the claims
where τ is a minimum evidence threshold determined by clinician review.

Legal Document Analysis with Statute Citations
Large language models (LLMs) applied to legal document analysis must not only extract relevant information but also provide verifiable citations to statutes, case law, and regulatory texts. This requires a combination of dense retrieval, semantic parsing, and hierarchical attention mechanisms to ensure precise grounding in authoritative sources.
Architecture for Statute-Aware Legal LLMs
The model architecture for legal document analysis typically integrates three key components:
- Legal Entity Recognition (LER): Identifies references to statutes, regulations, and case law using conditional random fields (CRFs) with domain-specific feature engineering.
- Cross-Document Attention: Computes attention scores between query text and potential citations using a modified version of the transformer's scaled dot-product attention:
where M is a legal citation mask that prioritizes statutory references over general text.
- Verification Layer: Validates extracted citations against a structured legal corpus using approximate nearest neighbor search in embedding space.
Training with Legal Citation Objectives
The model is trained with a multi-task objective combining:
where span loss identifies relevant text passages, citation loss maximizes precision in statute references, and consistency loss ensures alignment between generated explanations and cited authorities.
Hierarchical Statute Embeddings
Legal texts require specialized embedding approaches due to their nested structure (title → chapter → section → subsection). The embedding function E(s) for statute s combines:
where the graph neural network (GNN) processes the hierarchical relationships between legal provisions.
Evaluation Metrics for Legal QA Systems
Performance is measured using:
- Citation Precision (CP): Percentage of generated citations that are legally relevant to the query
- Statute Recall (SR): Proportion of applicable statutes correctly identified
- Legal Coherence Score (LCS): Expert-rated alignment between analysis and cited authorities
Benchmark results on the LexGLUE dataset show current state-of-the-art models achieve CP=0.82 and SR=0.76 when using hybrid retrieval-generation architectures.
Implementation Challenges
Key technical hurdles include:
- Handling conflicting or superseded statutes through temporal modeling
- Managing jurisdiction-specific variations in legal interpretation
- Balancing verbatim citation with paraphrased explanations
The most effective solutions employ dynamic knowledge graphs that track statute modifications and jurisdictional hierarchies, updated through continuous learning from official gazettes and court rulings.
4.3 Fact-Checking Assistants for Journalism
Modern journalism faces increasing challenges in verifying claims due to the rapid dissemination of information across digital platforms. Fact-checking assistants powered by explainable large language models (LLMs) address this by automating claim verification while providing transparent source attribution. These systems integrate three core components: retrieval-augmented generation (RAG), source reliability scoring, and claim-evidence alignment metrics.
Architecture of a Fact-Checking Pipeline
The pipeline begins with claim decomposition, where complex statements are broken into verifiable atomic propositions. For each proposition, the system performs:
- Multi-source retrieval: Queries diverse databases (news archives, scientific papers, government reports) using dense vector similarity over encoded claims.
- Stance detection: Computes the agreement between sources using a fine-tuned transformer model:
where s is the source text, c is the claim, and Wφ is a learned projection layer.
Source Reliability Estimation
Each retrieved document receives a credibility score combining:
- Provenance features: Publisher reputation, author expertise, and citation history
- Temporal decay: Exponential weighting of newer evidence
- Cross-validation: Agreement with other high-reliability sources
Explainability Mechanisms
The system generates human-interpretable justifications by:
- Highlighting conflicting evidence when sources disagree beyond a threshold τ
- Visualizing provenance chains showing how information propagated through sources
- Calculating confidence intervals for statistical claims using Monte Carlo sampling over source uncertainties
Case Study: Political Speech Analysis
When verifying a claim like "Country X has the highest tax burden in Europe," the system:
- Retrieves OECD tax reports, national budgets, and economic analyses
- Identifies that 2018-2022 data places Country X at 42.1% GDP vs. Denmark's 45.9%
- Flags the original claim as misleading with 87% confidence
- Outputs citable excerpts from primary sources with timestamps
Implementation Challenges
Key technical hurdles include:
- Latency constraints: Real-time verification requires optimized retrieval over compressed document indexes
- Multilingual verification: Cross-lingual embeddings must handle nuanced translations of claims
- Adversarial robustness: Detection of subtly manipulated sources designed to deceive automated systems

5. Key Research Papers in Explainable NLP
5.1 Key Research Papers in Explainable NLP
- Explainable AI (XAI): A systematic meta-survey of current challenges ... — Explainable artificial intelligence (XAI) has been proposed as a solution that can help to move towards more transparent AI and thus avoid limiting the adoption of AI in critical domains [1], [2].Generally speaking, according to [3], XAI focuses on developing explainable techniques that empower end-users in comprehending, trusting, and efficiently managing the new age of AI systems.
- Explainable AI Reloaded: Challenging the XAI Status Quo in the Era of ... — The field of Explainable AI (XAI) is concerned with developing techniques, concepts, and processes that can help stakeholders understand the reasons behind the AI system's decision-making [21, 34].For our purposes, we adopt a design lens in XAI that is sociotechnically-informed [12, 19, 34] and adopt the broad definition that an explanation is an answer to a why-question [11, 30, 35].
- A systematic review of Explainable Artificial Intelligence models and ... — AI was developed around 1950 in the computer science sector, and it copied the human mind to develop machines that can process, methodise, and perform based on the data given to the system, which will be useful when large amounts of datasets are used [1].AI machineries widely being used in the industrial domain and prompted to do a more research works in engineering fields such as NLP (natural ...
- Explainable Natural Language Processing - MIT Press — Explainable Natural Language Processing (NLP) is an emerging field, which has received significant attention from the NLP community in the last few years. At its core is the need to explain the predictions of machine learning models, now more frequently deployed and used in sensitive areas such as healthcare and law. The rapid developments in the area of explainable NLP have led to somewhat ...
- wangyongjie-ntu/Awesome-explainable-AI - GitHub — Benchmarking and Survey of Explanation Methods for Black Box Models, DMKD 2023. Post-hoc Interpretability for Neural NLP: A Survey, ACM Computing Survey 2022. Connecting the Dots in Trustworthy Artificial Intelligence: From AI Principles, Ethics, and Key Requirements to Responsible AI Systems and Regulation, Arxiv 2023. Explainable Artificial Intelligence (XAI): What we know and what is left ...
- Notions of explainability and evaluation approaches for explainable ... — The number of scientific articles, conferences and symposia around the world in eXplainable Artificial Intelligence (XAI) have significantly increased over the last decade [1], [2].This has led to the development of a plethora of domain-dependent and context-specific methods for dealing with the interpretation of machine learning (ML) models and the formation of explanations for humans.
- Title: Explainable AI: current status and future directions — Explainable Artificial Intelligence (XAI) is an emerging area of research in the field of Artificial Intelligence (AI). XAI can explain how AI obtained a particular solution (e.g., classification or object detection) and can also answer other "wh" questions. This explainability is not possible in traditional AI. Explainability is essential for critical applications, such as defense, health ...
- Explainability for Large Language Models: A Survey — Explainability 1 refers to the ability to explain or present the behavior of models in human-understandable terms [Doshi-Velez and Kim 2017; Du et al. 2019a].Improving the explainability of LLMs is crucial for two key reasons. First, for general end users, explainability builds appropriate trust by elucidating the reasoning mechanism behind model predictions in an understandable manner ...
- Survey on Explainable AI: From Approaches, Limitations and ... - Springer — Source-oriented (source-oriented (SO)) the sources that support building explanations can be either subjective (S) or objective (O) cognition, depending on whether the explanations are provided based on the fact or human experience. For example, in the medical field, if the explanation of a diagnosis is provided based on the patient's clinical symptoms and explains the cause and pathology in ...
- Explainable Artificial Intelligence: Importance, Use Domains, Stages ... — Therefore, Explainable Artificial Intelligence (XAI) has emerged to address the need for transparency in AI frameworks and to lower barriers to the widespread application of AI in important sectors. XAI is an approach to develop open techniques that let consumers comprehend and trust the evolving AI systems while being able to govern them successfully [].
5.2 Open-Source Implementations
- Studying LLM Performance on Closed- and Open-source Data - arXiv.org — Large language models (LLMs), like the ones used by CoPilot (Ziegler et al., 2022) are generally trained on very large source code corpora. LLMs are used in a wide variety of tasks, most notably for code-completion. Published research suggests that code-completion tools based on these models (Murali et al., 2023; Ziegler et al., 2022; Tabachnyk and Nikolov, ) offer substantial improvements in ...
- Evaluating the Efficacy of Open-Source LLMs in Enterprise-Specific RAG ... — This study examines various open-source LLMs, explores their in-tegration into RAG frameworks using enterprise-specific data, and assesses the performance of different open-source embeddings in enhancing the retrieval and generation process. Our findings indi-cate that open-source LLMs, combined with effective embedding
- A developer's guide to open source LLMs and generative AI — The future of open source LLMs. There's been a scurry of activity in the open source LLM world. "Developers are very active on some of these open source models," Aftandilian says. "They can optimize performance, explore new use cases, and push for new algorithms and more efficient data." And that's just the start.
- Closing the gap between open source and commercial large language ... — To quantitively assess fine-tuning technologies to enhance open-source LLMs for medical evidence summarization, we experimented with three broadly used open-sourced LLMs: PRIMERA 15, LongT5 14 ...
- Evaluation of open and closed-source LLMs for low-resource language ... — Despite the limited number of pre-trained models exclusively on Bengali, we assess the performance of six prominent LLMs, i.e., three closed-source (GPT-3.5, GPT-4o, Gemini) and three open-source (Aya 101, BLOOM, LLaMA) across key natural language processing (NLP) tasks, including text classification, sentiment analysis, summarization, and ...
- wangyongjie-ntu/Awesome-explainable-AI - GitHub — Pitfalls of Explainable ML: An Industry Perspective, Arxiv preprint 2021. Explainable Machine Learning in Deployment, FAT 2020. The elephant in the interpretability room: Why use attention as explanation when we have saliency methods, EMNLP Workshop 2020. A Survey of the State of Explainable AI for Natural Language Processing, AACL-IJCNLP 2020
- Open-Source vs Closed-Source LMS: Understanding Key LMS Technologies — Pablo Borbón Senior Director of Growth at Open LMS. Pablo Borbón is the Sr. Director of Strategic Operations for Open LMS, with a track record of 15 years in the EdTech industry. An electronic engineer, ecommerce specialist, and MBA candidate, Pablo brings a wealth of experience from both sides of education, serving as a university instructor in Colombia and excelling as a digital education ...
- 8 Top Open-Source LLMs for 2024 and Their Uses - DataCamp — The current generative AI revolution wouldn't be possible without the so-called large language models (LLMs). Based on transformers, a powerful neural architecture, LLMs are AI systems used to model and process human language.They are called "large" because they have hundreds of millions or even billions of parameters, which are pre-trained using a massive corpus of text data.
- A Comprehensive Guide to Explainable AI: From Classical Models to LLMs — The rise of artificial intelligence, particularly deep learning, has introduced remarkable advancements across numerous fields [27, 28].However, with these advancements comes a critical issue: the 'Black Box' problem [2].Many AI models, especially complex ones like neural networks and large language models (LLMs), are often regarded as black boxes due to their opaque decision-making ...
- Explainability for Large Language Models: A Survey — Explainability 1 refers to the ability to explain or present the behavior of models in human-understandable terms [Doshi-Velez and Kim 2017; Du et al. 2019a].Improving the explainability of LLMs is crucial for two key reasons. First, for general end users, explainability builds appropriate trust by elucidating the reasoning mechanism behind model predictions in an understandable manner ...
5.3 Recommended Courses and Tutorials
- Large Language Models (LLMs): A Comprehensive Guide — Open Source LLMs-Meta kicked off the trend of open source LLMs by releasing LLaMA for everyone earlier in 2023. Since then, a lot of open source LLMs have been released for commercial and research purposes such as Falcon, Mistral, and Zephyr to name a few. Here's a list of some of these LLMs. Papers —
- Understanding LLMs: A Comprehensive Overview from Training to Inference — Training LLMs require vast amounts of text data, and the quality of this data significantly impacts LLM performance. Pre-training on large-scale corpora provides LLMs with a fundamental understanding of language and some generative capability. ... Researchers can choose from these open-source LLMs to deploy applications that best suit their ...
- HITsz-TMG/awesome-llm-attributions - GitHub — [2023/09] Retrieving Evidence from EHRs with LLMs: Possibilities and Challenges Hiba Ahsan et al. arXiv. [2023/10] Learning to Plan and Generate Text with Citations Annoymous et al. OpenReview, ICLR 2024 [2023/10] 1-PAGER: One Pass Answer Generation and Evidence Retrieval Palak Jain et al. arxiv
- A comprehensive review of large language models: issues and solutions ... — A significant advancement in artificial intelligence is the development of large language models (LLMs). Despite opposition and explicit bans by some authorities, LLMs continue to play a transformative role, particularly in education, by improving language understanding and generation capabilities. This study explores LLMs' types, history, and training processes, alongside their application ...
- Online Courses - Learn Anything, On Your Schedule | Udemy — Udemy is an online learning and teaching marketplace with over 250,000 courses and 73 million students. Learn programming, marketing, data science and more. Search bar. Site navigation Explore by Goal. Launch a new career. Certification preparation.
- Source-Aware Training Enables Knowledge Attribution in Language Models — To this end, we explore source-aware training to enable an LLM to cite the pretraining source supporting its parametric knowledge. Our motivation is three-fold. First, a significant portion of an LLM's knowledge is acquired during pretraining, therefore citing evidence for this parametric knowledge can greatly enhance the LLM trustworthiness.
- How Can We Use LLMs for EDM Tasks? The Case of Course ... - ResearchGate — We hope that this work can be a source of motivation for researchers to look into how LLMs might improve recommendation performance and other educational data mining tasks.
- Source-Aware Training Enables Knowledge Attribution in - arXiv.org — To this end, we explore source-aware training —a post-pretraining recipe that enables a LLM to cite its pretraining data based on its parametric knowledge. Our motivation is three-fold. First, a significant portion of an LLM's knowledge is acquired during pretraining, therefore citing evidence for this parametric knowledge can greatly enhance the LLM trustworthiness.
- What Should Data Science Education Do With Large Language Models? — 1. Introduction. The rapid advancements in artificial intelligence have led to the development of powerful tools, one of the most notable being large language models (LLMs) such as ChatGPT by OpenAI (Brown et al., 2020; OpenAI, 2023).These models have demonstrated remarkable capabilities in understanding and generating humanlike text, often outperforming traditional algorithms in various ...








