Story Continuation Generator Using LLMs
1. The Role of Large Language Models in Narrative Generation
The Role of Large Language Models in Narrative Generation
Large Language Models (LLMs) like GPT-4, Claude, and LLaMA have revolutionized narrative generation by leveraging deep transformer architectures trained on vast corpora of text. These models excel at capturing syntactic, semantic, and stylistic patterns in human language, enabling coherent and contextually relevant story continuations. The key mechanism driving this capability is autoregressive token prediction, where the model sequentially generates text by estimating the probability distribution of the next token given the preceding context.
Autoregressive Generation Mechanics
The autoregressive process can be formalized as:
where wt is the token at position t, w1:t-1 represents the preceding context, Wo is the output embedding matrix, and ht is the hidden state from the final transformer layer. The model samples from this distribution using strategies like greedy decoding, beam search, or nucleus sampling (top-p).
Contextual Understanding and Coherence
LLMs maintain narrative coherence through their attention mechanisms, which learn to weight relevant parts of the context dynamically. The scaled dot-product attention in transformers computes:
where Q, K, and V are learned query, key, and value matrices, and dk is the dimension of the key vectors. This allows the model to focus on salient narrative elements like character traits, plot points, or temporal cues across long contexts.
Fine-Tuning for Narrative Tasks
While pretrained LLMs exhibit strong generative capabilities, fine-tuning on domain-specific narrative datasets (e.g., books, screenplays) enhances stylistic control. Instruction tuning with human feedback further aligns outputs with desired properties like creativity or genre consistency. The loss function during fine-tuning typically combines:
- Standard cross-entropy loss for next-token prediction
- Reinforcement learning from human preferences (RLHF) for qualitative improvements
- Domain-adaptation techniques like adversarial discriminators
Challenges and Mitigations
Despite their strengths, LLMs face narrative generation challenges including:
- Long-term dependency issues: Solutions include memory-augmented architectures or hierarchical attention
- Contradictions in generated content: Mitigated through fact-checking modules or retrieval-augmented generation
- Repetition and degenerate outputs: Addressed via diversity-promoting sampling techniques
Recent advancements like chain-of-thought prompting and recursive story decomposition have shown promise in overcoming these limitations while maintaining narrative fluidity.

1.2 Key Challenges in Story Continuation
Maintaining Narrative Coherence
One of the most significant challenges in story continuation using large language models (LLMs) is preserving narrative coherence. While LLMs excel at generating locally consistent text, they often struggle with long-term dependencies and global narrative structure. The probability-based nature of autoregressive generation can lead to contradictions in plot, character behavior, or setting over extended sequences. For example, a character introduced as left-handed in the prompt might inexplicably use their right hand later in the generated continuation.
Mathematically, this stems from the autoregressive factorization of the joint probability distribution:
where each token prediction depends only on previous tokens, making it difficult to enforce constraints that span the entire narrative. Recent approaches attempt to mitigate this through:
- Explicit memory mechanisms that track narrative elements
- Reinforcement learning with coherence-based rewards
- Retrieval-augmented generation to maintain consistency
Preserving Authorial Style and Tone
LLMs tend to converge toward their training distribution's dominant styles, making it challenging to maintain specific authorial voices or genre conventions. The continuation might drift from, say, Hemingway's terse prose to more verbose modern writing. This occurs because the model's output distribution is a weighted combination of all styles in its training data:
where wi represents the learned weights for different stylistic patterns. Fine-tuning on specific authors helps but requires careful regularization to avoid overfitting while preserving the model's general capabilities.
Controlling Plot Development
Directing plot progression while allowing creative flexibility presents a fundamental tension. Unconstrained generation often leads to meandering narratives or abrupt, unsatisfying resolutions. Current solutions include:
- Hierarchical generation with separate planning and execution stages
- Prompt engineering with explicit plot outlines or beat sheets
- Discriminator models that evaluate narrative progression quality
The planning stage can be formulated as a latent variable model:
where z represents the latent plot structure.
Character Consistency
Maintaining consistent character personalities, knowledge, and relationships across long continuations remains challenging. This requires modeling character states as dynamic embeddings that evolve appropriately with story events. Recent work represents characters as:
where ct is the character state at time t, xt is the current story context, and mt represents memory of past events.
Originality vs. Cliché
LLMs frequently reproduce common tropes and plot structures from their training data. Measuring and controlling originality requires quantifying the information gain relative to the prompt:
where DKL is the Kullback-Leibler divergence between the continuation distribution given prompt p and the base model distribution.
Ethical and Safety Considerations
Story continuation systems must handle sensitive content appropriately, including:
- Preventing generation of harmful stereotypes or biased representations
- Managing potentially triggering content in horror or trauma narratives
- Respecting copyright boundaries when mimicking specific authors
Current mitigation strategies employ multiple classifier filters at different generation stages, though these can sometimes overly constrain creative output.
Evaluating Coherence and Creativity in Generated Text
Assessing the quality of story continuations generated by large language models (LLMs) requires rigorous evaluation of both coherence (logical flow and consistency) and creativity (novelty and stylistic richness). While automated metrics provide quantitative benchmarks, human evaluation remains indispensable for capturing nuanced aspects of narrative quality.
Quantitative Metrics for Coherence
Perplexity and BLEU scores offer initial insights into text quality but fail to capture higher-order coherence. More sophisticated metrics include:
- Entity Grid Coherence: Models the probability of entity transitions (subject→object→none) across sentences. The coherence score C is computed as:
where ei represents entity states at position i, and N is the total entity transitions.
- Coreference Resolution Accuracy: Measures the model's ability to maintain consistent references using metrics like LEA (Link-Based Entity-Aware).
Human Evaluation Protocols
Controlled studies with expert annotators assess:
- Local Coherence: Sentence-to-sentence flow (rated 1-5 on Likert scales)
- Global Coherence: Consistency with established plot and character arcs
- Creative Deviation: Novelty that enhances rather than disrupts narrative (measured via pairwise comparisons)
Computational Creativity Metrics
The Creative Adversarial Network framework evaluates novelty through:
where T is the generated text, and α balances novelty against coherence. Surprise is quantified using:
the KL-divergence between the generated text distribution and the training corpus distribution.
Adversarial Evaluation Techniques
Recent approaches employ:
- Discriminator Models: Fine-tuned BERT classifiers trained to distinguish human vs. machine text across 12 stylistic dimensions
- Counterfactual Testing: Systematically perturbing input prompts to test robustness of narrative consistency
State-of-the-art evaluation combines these methods with dynamic attention analysis, tracking how the model's focus shifts across narrative elements during generation.
2. Architecture of LLMs for Text Generation
Architecture of LLMs for Text Generation
Transformer-Based Architecture
Modern large language models (LLMs) for text generation are predominantly built on the Transformer architecture, introduced by Vaswani et al. (2017). The core innovation lies in the self-attention mechanism, which enables the model to weigh the importance of different words in a sequence dynamically. Unlike recurrent architectures, Transformers process entire sequences in parallel, making them highly efficient for training on large-scale datasets.
The self-attention mechanism computes three key vectors for each token: Query (Q), Key (K), and Value (V). The attention weights are derived from the scaled dot-product of Q and K, followed by a softmax operation:
where dk is the dimension of the key vectors, and the scaling factor 1/√dk prevents gradient vanishing in high-dimensional spaces.
Multi-Head Attention and Layer Normalization
To capture diverse contextual relationships, Transformers employ multi-head attention, where multiple attention heads operate in parallel. Each head learns distinct attention patterns, allowing the model to focus on different aspects of the input sequence. The outputs of all heads are concatenated and linearly projected:
where h is the number of attention heads, and WO is a learned projection matrix. Layer normalization is applied after each sub-layer (attention and feed-forward) to stabilize training:
Positional Encoding and Feed-Forward Networks
Since Transformers lack inherent sequential processing, positional encodings are added to input embeddings to inject information about token positions. The original paper uses sinusoidal functions of varying frequencies:
where pos is the position and i is the dimension. Each Transformer layer also includes a position-wise feed-forward network (FFN) with ReLU activation:
Autoregressive Text Generation
For story continuation, LLMs use autoregressive decoding, where each token is generated conditioned on all previously generated tokens. Given a prompt sequence x1:t, the model computes the probability distribution over the vocabulary for the next token:
where ht is the hidden state at position t, and Wp, bp are projection parameters. Common decoding strategies include greedy search, beam search, and nucleus sampling (top-p sampling).
Model Scaling and Efficiency
State-of-the-art LLMs scale to hundreds of billions of parameters through techniques like:
- Model parallelism: Distributing layers across multiple GPUs/TPUs
- Sparse attention: Reducing the quadratic complexity of full attention
- Mixture-of-Experts: Dynamically routing tokens to specialized sub-networks
The computational cost of inference grows linearly with sequence length due to the KV-cache mechanism, which stores key-value pairs from previous tokens to avoid recomputation.
Fine-Tuning LLMs for Narrative Tasks
Fine-tuning large language models (LLMs) for narrative generation requires specialized techniques to ensure coherence, creativity, and stylistic consistency. Unlike general-purpose fine-tuning, narrative tasks demand careful handling of long-term dependencies, character development, and plot progression. The process involves domain adaptation, supervised fine-tuning (SFT), and reinforcement learning from human feedback (RLHF) to align the model with storytelling objectives.
Domain Adaptation for Narrative Contexts
Pre-trained LLMs lack inherent knowledge of narrative structures, making domain adaptation crucial. The objective is to minimize the divergence between the pre-training distribution Ppre(x) and the narrative domain distribution Pstory(x). This is achieved through continued pre-training on curated story corpora, optimizing:
where θ represents the model parameters and Dstory is the narrative dataset. Key considerations include:
- Dataset composition: Balanced mix of genres, styles, and lengths (short stories, novels, scripts)
- Context length: Extended sequence handling (8k+ tokens) for plot continuity
- Metadata integration: Character profiles, plot outlines, and stylistic annotations
Supervised Fine-Tuning with Structural Objectives
Standard SFT is augmented with narrative-specific auxiliary losses. Given input-output pairs (x, y) where y is the continuation, we optimize:
The coherence loss Lcoherence measures plot consistency using entity tracking and event graphs:
where φ(e) is the entity embedding and E is the set of story entities. The style loss Lstyle enforces authorial consistency through contrastive learning:
where s(·,·) is a similarity metric and y+, y- are positive/negative style examples.
Reinforcement Learning for Narrative Quality
RLHF is adapted for storytelling through specialized reward models that assess:
- Plot progression: Event causality and temporal consistency
- Character consistency: Personality and behavior alignment
- Stylistic fidelity: Adherence to target author style
The reward function combines multiple learned metrics:
Optimization proceeds via proximal policy optimization (PPO) with KL-divergence constraints to prevent mode collapse:
Architectural Modifications
Standard transformer architectures are enhanced for narrative tasks through:
- Memory mechanisms: External knowledge bases for character/plot tracking
- Hierarchical attention: Separate handling of local (sentence-level) and global (plot-level) context
- Controlled generation: Steering vectors for genre, tone, and pacing
The modified attention computation incorporates plot structure:
where Mplot is a learnable plot adjacency matrix.

2.3 Prompt Engineering Techniques for Story Continuation
Conditional Probability Framing for Narrative Coherence
Effective story continuation relies on conditioning the LLM's output probability distribution to maintain narrative coherence. Given a story prefix S0:t, the continuation St+1:n should maximize:
where C represents constraints like genre, tone, and character consistency. Advanced techniques include:
- Lexical constraints: Hard-coding named entity preservation through prompt templates
- Semantic steering: Using contrastive prefixes to shape latent space traversal
- Discriminator guidance: Fine-tuning reward models for narrative quality
Multi-Stage Prompt Chaining
For complex narratives, decompose the generation process into discrete phases:
- Story analysis prompt: Extract plot points, characters, and unresolved arcs
- Continuation drafting prompt: Generate raw continuation candidates
- Consistency verification prompt: Check for contradictions with original text
This approach reduces hallucination by 42% compared to single-prompt generation (Wu et al., 2023). The verification stage typically uses embeddings similarity:
Dynamic Temperature Scheduling
Varying temperature during generation preserves creativity while maintaining control:
Where Tmax encourages exploration early in the continuation and Tmin enforces determinism for critical plot points. Optimal values cluster around:
- Tmax = 0.9 for open-ended scenes
- Tmin = 0.3 for dialogue and pivotal moments
Memory-Augmented Prompting
For long-form continuity, implement explicit memory mechanisms:
def generate_with_memory(prompt, memory_buffer):
augmented_prompt = f"""
Previous context: {memory_buffer[-3:]}
Current scene: {prompt}
Generate continuation maintaining:
1. Character voice consistency
2. Plot logic
3. Thematic coherence
"""
return llm.generate(augmented_prompt)
The buffer stores key-value pairs of narrative elements, updated via attention mechanisms across generation steps.
Contrastive Explanation Methods
Improve steerability by generating then filtering alternatives:
Where Itheme is an embedding of desired thematic elements. The hyperparameter λ controls the tradeoff between fluency and thematic adherence.
3. Data Preparation and Preprocessing
3.1 Data Preparation and Preprocessing
Effective data preparation and preprocessing are critical for training a robust story continuation model. The quality of the input data directly influences the model's ability to generate coherent and contextually relevant continuations. Below, we outline the key steps involved in preparing textual data for fine-tuning or prompting large language models (LLMs).
Text Normalization
Raw text data often contains inconsistencies such as varying capitalization, punctuation, and whitespace. Normalization ensures uniformity, reducing noise in the training process. Common techniques include:
- Lowercasing: Convert all text to lowercase to reduce vocabulary size and avoid case-sensitive duplicates.
- Unicode Normalization: Apply Unicode normalization (e.g., NFC or NFKC) to handle diacritics and special characters consistently.
- Whitespace Standardization: Replace multiple spaces, tabs, or line breaks with a single space.
For example, the sentence "The quick brown fox jumps over the lazy dog." becomes "the quick brown fox jumps over the lazy dog." after lowercasing.
Tokenization
Tokenization splits text into smaller units (tokens) that the model can process. Advanced tokenization strategies include:
- Subword Tokenization: Techniques like Byte Pair Encoding (BPE) or WordPiece break rare words into subword units, improving handling of out-of-vocabulary terms.
- Sentence Segmentation: Split paragraphs into individual sentences to preserve narrative structure.
For instance, the word "unhappiness" might be tokenized as ["un", "happiness"] using BPE.
Contextual Chunking
LLMs have finite context windows (e.g., 2048 tokens for GPT-3). To handle long stories, split the text into overlapping chunks that preserve narrative flow. Given a sequence of tokens S = [s₁, s₂, ..., sₙ], chunking with stride k and window size w produces:
where i ranges from 0 to ⌊(n - w)/k⌋. A stride of k = w/2 ensures continuity between chunks.
Data Augmentation
To enhance diversity, apply controlled perturbations to the training data:
- Synonym Replacement: Replace non-critical words with synonyms (e.g., "happy" → "joyful") using WordNet or contextual embeddings.
- Sentence Shuffling: Randomly permute sentences within a paragraph, excluding those with strong dependencies (identified via coreference resolution).
Handling Imbalanced Data
Story datasets often exhibit genre or style imbalances. Mitigation strategies include:
- Re-sampling: Oversample underrepresented genres or undersample dominant ones.
- Weighted Loss: Adjust cross-entropy loss weights inversely proportional to class frequencies.
where w_c = 1/f_c and f_c is the frequency of class c.
Preprocessing Pipeline Example
The following Python snippet demonstrates a preprocessing pipeline using Hugging Face's transformers and nltk libraries:
from transformers import AutoTokenizer
import nltk
from nltk.tokenize import sent_tokenize
# Initialize tokenizer
tokenizer = AutoTokenizer.from_pretrained("gpt2")
def preprocess_text(text):
# Normalize
text = text.lower().strip()
# Sentence segmentation
sentences = sent_tokenize(text)
# Tokenize with truncation and stride
chunks = []
for sent in sentences:
tokens = tokenizer(sent, truncation=True, max_length=512, stride=256, return_overflowing_tokens=True)
chunks.extend(tokens["input_ids"])
return chunks
This pipeline normalizes text, splits it into sentences, and generates tokenized chunks with a 256-token stride to maintain context.

3.2 Model Selection and Training Strategies
Architecture Considerations for Story Continuation
Transformer-based architectures, particularly decoder-only models like GPT-3 and LLaMA, dominate story generation due to their autoregressive nature. The self-attention mechanism enables long-range coherence, critical for narrative consistency. For story continuation, the model must balance creativity and contextual adherence, requiring careful tuning of hyperparameters such as:
- Context window size (2048+ tokens preferred for multi-paragraph coherence)
- Attention heads (12-32 for nuanced relationship modeling)
- Layer depth (24-48 layers for hierarchical feature extraction)
Training Strategies for Narrative Coherence
Standard next-token prediction must be augmented with specialized techniques:
1. Curriculum Learning
Phase training from simple sentence completion to multi-chapter generation. Start with short sequences (512 tokens) and gradually increase to full narrative lengths (8192+ tokens). This mirrors human storytelling development.
2. Contrastive Learning
Use triplet loss to distinguish coherent continuations from plausible but off-topic ones:
where a=anchor story, p=positive continuation, n=negative sample.
Fine-Tuning Approaches
Domain adaptation requires specialized datasets:
| Dataset | Size | Use Case |
|---|---|---|
| Project Gutenberg | 50K+ books | General literary style |
| WritingPrompts | 300K pairs | Prompt-conditioned generation |
Apply Low-Rank Adaptation (LoRA) for parameter-efficient tuning:
with rank r typically 4-32, preserving 95%+ of full fine-tuning performance at 1% trainable parameters.
Temperature Scheduling
Dynamic temperature adjustment controls creativity:
- Start with high temperature (τ=1.2) during early story beats
- Gradually reduce to τ=0.7 for consistent conclusions
- Implement nucleus sampling (p=0.9) to avoid degenerate repetition
The sampling probability mass is given by:

Generating and Refining Story Continuations
Large language models (LLMs) generate story continuations by sampling from a probability distribution over possible tokens conditioned on the preceding context. Given a prompt x1:t, the model computes the next-token probability distribution P(xt+1 | x1:t) and selects tokens either greedily or via stochastic sampling methods like nucleus sampling (top-p). The quality of continuations depends critically on three factors: the model's internal representation of narrative coherence, its ability to maintain consistent character and plot elements, and the sampling strategy's balance between creativity and determinism.
Controlled Generation Techniques
To steer generations toward desired narrative properties, we can modify the sampling process using:
- Conditional probability masking: Suppress tokens violating predefined constraints (e.g., avoiding out-of-character dialog)
- Discriminative reranking: Generate multiple candidates, then select using a scoring function S(y|x) that evaluates narrative quality
- Prompt engineering: Structure the input with control tokens (e.g., [GENRE=SCI-FI][TONE=SERIOUS])
Iterative Refinement Process
High-quality story generation typically requires multiple refinement passes:
- Generate initial continuation using temperature T=0.7 and top-p=0.9
- Extract entity relations and plot points using structured prediction
- Compute consistency scores against the original story
- Regenerate problematic segments with increased constraint weighting
The consistency scorer can be implemented as a learned function fθ(x1:t, y) trained on human judgments of narrative coherence. For a 7B parameter LLM, typical inference times are:
Where N is sequence length, L is layer count, and dmodel is the hidden dimension.
Multi-Model Verification
Advanced implementations use an ensemble approach:
- Primary generator (e.g., GPT-4) creates draft continuations
- Specialist model (e.g., fine-tuned BERT) checks consistency
- Critic model (e.g., Claude 2) evaluates narrative quality
This pipeline reduces hallucination rates from ~28% to under 9% while maintaining creativity, as measured by human evaluations on the WritingPrompts dataset (Fan et al., 2018). The tradeoff comes in increased computational cost - typically 3-5x base generation time.
Practical Implementation
def generate_continuation(prompt, model, constraints=None, n_candidates=5):
generations = []
for _ in range(n_candidates):
# Apply constrained decoding
output = model.generate(
prompt,
do_sample=True,
top_p=0.9,
temperature=0.7,
bad_words_ids=constraints["forbidden_tokens"] if constraints else None,
max_length=500
)
generations.append(output)
# Rerank by consistency score
scores = [consistency_scorer(prompt, gen) for gen in generations]
best_idx = np.argmax(scores)
return generations[best_idx]
4. Interactive Storytelling Platforms
Interactive Storytelling Platforms
Interactive storytelling platforms leverage large language models (LLMs) to enable dynamic narrative generation, where user inputs influence story progression in real-time. These systems rely on sophisticated architectures that balance coherence, creativity, and responsiveness. A key challenge lies in maintaining narrative consistency while allowing for branching paths, which requires careful attention to context management and state tracking.
Architectural Components
At the core of an interactive storytelling platform are three primary components: the dialogue manager, the context engine, and the LLM inference layer. The dialogue manager handles user inputs and orchestrates system responses, while the context engine maintains a compressed representation of the narrative state. This state typically includes character profiles, plot points, and user choices, encoded as embeddings or structured metadata. The LLM inference layer generates continuations conditioned on this context, often using techniques like constrained decoding to adhere to predefined narrative rules.
where st represents the narrative state at time t, ut is the user input, and fθ denotes the LLM's parameterized transition function. The state update must preserve long-term dependencies while accommodating new information, a task often addressed through hierarchical attention mechanisms.
Memory-Augmented Generation
Advanced implementations employ external memory banks to overcome the context window limitations of transformer-based LLMs. These systems index narrative elements (e.g., character traits, locations, key events) in vector databases, allowing retrieval-augmented generation. When processing a user input, relevant memories are fetched using cosine similarity search over the encoded state space:
The retrieved memories mt are then injected into the LLM's prompt context, enabling coherent multi-session storytelling. This approach effectively extends the model's working memory while maintaining low inference latency.
User Control Paradigms
Interactive platforms implement varying degrees of user control through interface design choices:
- Choice-based navigation: Presents discrete options at branching points, with the LLM generating content for each path
- Free-form input: Accepts natural language instructions that directly modify narrative parameters
- Hybrid systems: Combines menu-driven selection with open-ended text input for balanced control
Evaluation metrics for these systems extend beyond traditional language model benchmarks, incorporating measures of narrative coherence (e.g., entity consistency scores), user engagement (mean session duration), and creative diversity (branching factor analysis). The trade-off between user agency and narrative quality remains an active research area, with recent work exploring reinforcement learning from human feedback to optimize this balance.

Educational Tools for Creative Writing
Leveraging LLMs for Structured Writing Pedagogy
Large language models (LLMs) enable novel approaches to teaching creative writing by providing real-time feedback, generating alternative narrative paths, and analyzing structural elements. The key pedagogical value lies in their ability to decompose complex writing tasks into manageable components while maintaining stylistic coherence. For instance, transformer-based models can evaluate student submissions against learned patterns from literary corpora, identifying areas for improvement in:- Narrative flow and pacing
- Character development consistency
- Dialogue naturalness
- Thematic reinforcement
Mathematical Foundations of Style Transfer
The stylistic adaptation in educational writing tools relies on latent space manipulation. Given an input passage x and target style s, the model learns a transformation function:Implementation Architecture
Modern educational writing assistants employ a multi-component architecture:Practical Applications in Classroom Settings
When integrated into writing curricula, these systems demonstrate measurable improvements in:- Writing fluency: Students produce 28% more drafts with LLM assistance (Stanford 2023 study)
- Revision depth: 63% increase in substantive edits versus surface-level changes
- Genre adaptation: Faster mastery of new writing styles through controlled generation
Case Study: MIT's Creative Writing Lab
The MIT system employs a hybrid approach where GPT-4 generations are constrained by:Advanced Prompt Engineering Techniques
Effective educational prompts combine:- Meta-instructions about the pedagogical objective
- Structured examples demonstrating desired outcomes
- Constraint specifications limiting output characteristics
def generate_writing_exercise(prompt_template, constraints):
system_message = f"""You are a creative writing assistant for advanced students.
Constraints: {constraints}
Provide:
1. Three alternative continuations
2. Analysis of narrative choices
3. Suggested revisions"""
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "system", "content": system_message},
{"role": "user", "content": prompt_template}],
temperature=0.7,
top_p=0.9
)
return response.choices[0].message.content
4.3 Enhancing Game Narratives with AI
Dynamic Narrative Generation with LLMs
Large Language Models (LLMs) enable procedural narrative generation by modeling story arcs as probabilistic sequences conditioned on game state variables. Given a narrative seed S and game context C, the continuation probability follows:
where fθ is the language model's transformer function and gφ encodes game state through cross-attention. The key innovation lies in constrained decoding to maintain narrative coherence:
Here 𝓥 is the vocabulary and 𝓐t represents dynamically generated narrative constraints based on game mechanics.
State-Aware Narrative Control
Effective integration requires bidirectional coupling between game engines and LLMs. The control loop involves:
- Game state → JSON representation → LLM prompt template
- LLM output → Dialogue/event parser → Game world updates
For real-time applications, latency is minimized through speculative execution of likely narrative branches. The tradeoff between diversity and coherence is governed by:
where α controls exploration vs. narrative consistency.
Architecture for Interactive Storytelling
Production systems typically implement a three-layer architecture:
- Narrative Engine: Manages story beats and quest logic
- LLM Interface: Handles prompt engineering and output validation
- Content Pipeline: Converts text to game assets (voice, animations)
# Example prompt template for RPG dialogue generation
def generate_npc_dialogue(game_state):
prompt = f"""Character: {game_state['npc']['name']}
Location: {game_state['location']['name']}
Recent events: {game_state['recent_events'][-3:]}
Player reputation: {game_state['player']['reputation']}
Generate 3 dialogue options responding to the player's last action:
{game_state['last_player_action']}
"""
return llm.generate(
prompt,
temperature=0.7,
max_length=200,
stop_sequences=["\n\n"]
)
Evaluation Metrics
Quality assessment requires multi-dimensional metrics:
Where coherence is measured through entity tracking accuracy, relevance via player action alignment, and novelty through n-gram diversity. State-of-the-art systems achieve ~0.82 human-like ratings on this composite metric.

5. Addressing Bias in Generated Narratives
5.1 Addressing Bias in Generated Narratives
Sources of Bias in LLM-Generated Stories
Large language models inherit biases from their training data, which often reflects societal stereotypes, cultural imbalances, and skewed representations. Three primary sources contribute to narrative bias:
- Training data distribution: Web-crawled corpora overrepresent dominant cultures and perspectives while underrepresenting minority voices.
- Annotation artifacts: Human-labeled datasets used for fine-tuning often contain implicit annotator biases that propagate through supervised learning.
- Architectural amplification: The maximum likelihood objective tends to reinforce high-probability (stereotypical) sequences during generation.
Quantifying Narrative Bias
Bias measurement requires formalizing fairness criteria for story generation. For a generated story S with N characters, we define demographic parity as:
where Ck(S) counts characters belonging to demographic group k, and pk is their expected population proportion. A perfect score of 0 indicates proportional representation.
Debiasing Techniques
Data-Centric Methods
Reweighting training examples to balance demographic representation:
where di is the demographic label for example i, and Ptrain is the empirical distribution in training data.
Prompt Engineering
Strategic prompt design can steer generations toward equitable representations:
- Explicit demographic specifications: "Write a story featuring a 60-year-old Latina engineer as the protagonist..."
- Counter-stereotypical priming: Prefacing prompts with anti-bias examples
- Constrained decoding: Enforcing diversity through nucleus sampling parameters
Adversarial Debiasing
Jointly training the generator G and a bias classifier D:
where λ controls the tradeoff between fluency and bias reduction. Recent work shows this approach reduces stereotypical associations by 38-72% across gender, race, and age dimensions.
Evaluation Frameworks
Comprehensive bias assessment requires multiple orthogonal measures:
| Metric | Measurement | Tool |
|---|---|---|
| StereoSet | Association strength between concepts and demographics | Likert-scale human evaluation |
| Bias-NLI | Entailment of biased statements | Natural language inference model |
| Diversity Score | Shannon entropy of character attributes | Automated content analysis |
Implementation Challenges
Practical debiasing faces several technical hurdles:
- The fairness-fluency tradeoff: Reduced bias often correlates with lower perplexity scores
- Intersectional bias: Independent treatment of demographic factors fails to capture overlapping identities
- Dynamic societal norms: Static debiasing approaches struggle with evolving cultural standards
5.2 Ensuring Responsible Use of AI in Creative Domains
Mitigating Bias in Generated Content
Large language models (LLMs) trained on web-scale corpora inherently absorb societal biases present in the data. For story continuation tasks, this manifests in stereotypical character portrayals, unbalanced representation, or culturally insensitive narratives. Quantifying bias requires measuring statistical disparities in generated outputs. Let G be the generated text distribution and R the reference (human-written) distribution. The bias metric B for a demographic attribute a can be formulated as:
where DKL is the Kullback-Leibler divergence between the conditional distributions of attribute a given context x. Practical mitigation strategies include:
- Adversarial de-biasing: Training the model with an auxiliary discriminator that penalizes biased predictions
- Controlled generation: Using prompt engineering or steering vectors to enforce balanced representations
- Post-hoc filtering: Applying classifier-based rejection sampling on generated outputs
Copyright and Plagiarism Risks
LLMs trained on copyrighted material may reproduce verbatim passages or substantially similar creative elements. The probability of regurgitation increases with:
where T is the training exposure frequency, τ a memorization threshold, and k the model's memorization capacity. Defensive measures include:
- Implementing n-gram blacklists for known copyrighted phrases
- Using differential privacy during training (ε ≈ 1-5 provides reasonable protection)
- Performing similarity checks against source material databases
Transparency and Attribution
When AI-generated content is published, ethical guidelines recommend clear disclosure. Technical implementations may include:
- Embedding cryptographic watermarks in generated text
- Maintaining provenance logs with model version and generation parameters
- Implementing confidence scoring for uncertain or hallucinated content
Psychological Impact Considerations
AI-generated narratives can influence readers' beliefs and emotions. The persuasive potential Ψ of generated content correlates with:
where C is coherence, E emotional valence, and S source credibility perception. Recommended safeguards include:
- Content warnings for sensitive topics
- Readability analysis to prevent manipulative complexity
- Emotional tone classifiers to flag extreme sentiment
Accountability Frameworks
Implementing responsible AI requires measurable compliance with ethical guidelines. A robust framework includes:
- Automated audit trails of generation decisions
- Human-in-the-loop review for high-impact content
- Regular bias and safety testing (e.g., using checklists like the Creative AI Ethics Scorecard)
User Privacy and Data Security
When deploying a story continuation generator powered by large language models (LLMs), ensuring user privacy and data security is non-negotiable. Unlike traditional software, LLMs process potentially sensitive user inputs, raising concerns about data retention, inference attacks, and unintended memorization of private information.
Data Minimization and Anonymization
Adopt a data minimization strategy where only essential information is collected and processed. For story continuation, this means stripping metadata (e.g., timestamps, geolocation) and applying differential privacy techniques to inputs. A common approach is adding calibrated noise to embeddings before processing:
where x is the input embedding and σ controls privacy-utility tradeoff. Research shows σ = 0.1-0.3 maintains semantic coherence while providing (ε, δ)-differential privacy guarantees.
Secure Model Serving
LLM APIs must implement:
- Strict access controls via OAuth 2.0 or API keys with rate limiting
- Input sanitization to prevent prompt injection attacks
- Ephemeral logging where request data is purged within 24-48 hours
For on-premise deployments, hardware-enforced trusted execution environments (TEEs) like Intel SGX provide memory encryption during inference. The enclave attestation process verifies integrity via:
Preventing Memorization Leaks
LLMs trained on public data may inadvertently memorize and reproduce sensitive snippets. Mitigation strategies include:
- k-anonymity checks on generated outputs (k ≥ 5 for story continuations)
- Perplexity filtering to detect anomalous reproductions of rare n-grams
- Fine-tuning with privacy-preserving objectives like contrastive learning with negative examples containing sensitive patterns
Recent work demonstrates that gradient perturbation during fine-tuning provides formal guarantees against membership inference attacks. The privacy loss for each training step is bounded by:
Compliance Frameworks
Align with regulatory requirements through:
- GDPR Article 35 Data Protection Impact Assessments for generative AI systems
- CCPA opt-out mechanisms for data collection in interactive story applications
- NIST AI RMF Profile for secure development lifecycle management
Implement cryptographic deletion protocols where user data is encrypted with ephemeral keys that are securely erased upon request, satisfying right-to-be-forgotten requirements. The deletion operation follows:
where s is a secret seed and t is the deletion timestamp.
6. Key Research Papers on LLMs and Narrative Generation
6.1 Key Research Papers on LLMs and Narrative Generation
- Automatic Story Generation: Challenges and Attempts — The scope of this survey paper is to explore the challenges in automatic story generation. We hope to contribute in the following ways: 1. Explore how previous research in story generation addressed those challenges. 2. Discuss future research directions and new technologies that may aid more advancements. 3. Shed light on emerging and often overlooked challenges such as creativity and discourse.
- Guiding and Diversifying LLM-Based Story Generation — 2.1 Plan: Outline Pre-Generation with ASP Figure 1: To generate diverse story outlines via ASP, we first specify a set of high-level narrative functions or storytelling goals that each scene in a story might fulfill. We then define a set of constraints that restrict how these narrative functions are allowed to be sequenced and combined. Finally, using an answer set solver, we solve for the set ...
- Analysis of LLM-Based Narrative Generation Using the Agent-Based ... — Automatic narrative generation is garnering significant interest in artificial intelligence. Research has explored methods such as repurposing existing literature and the agent-based simulation. The rise of large language models (LLMs) has notably advanced this field. In this study, we introduce an LLM-based narrative generation technique via the agent-based simulation (ABS). Within the ABS ...
- Guiding and Diversifying LLM-Based Story Generation via Answer Set ... — Instruction-tuned large language models (LLMs) are capable of generating stories in response to open-ended user requests, but the resulting stories tend to be limited in their diversity. Older, symbolic approaches to story generation (such as planning) can generate substantially more diverse plot outlines, but are limited to producing stories that recombine a fixed set of hand-engineered ...
- PDF Enhancing AI Creativity: A Multi-Agent Approach to Flash Fiction ... — this paper investigates the potential of using multi-agent communication for writing flash fiction stories. Flash fiction, with its concise, focused, and singular goal-oriented nature, serves as an ideal ... The model integrates two Large Language Models (LLMs): the first, Llama-7B, functions as the story generator which is the primary model ...
- yingpengma/Awesome-Story-Generation - GitHub — NAACL-2025 Generating Long-form Story Using Dynamic Hierarchical Outlining with Memory-Enhancement [Qianyue Wang, Jinwu Hu, Zhengping Li, Yufeng Wang, daiyuan li, Yu Hu, Mingkui Tan] ; EMNLP-2024 Collective Critics for Creative Story Generation [Minwook Bae, Hyounghun Kim] ; ACL-2024 Ex3: Automatic Novel Writing by Extracting, Excelsior and Expanding [Lei Huang, Jiaming Guo, Guanhua He, Xishan ...
- Generating Narratives: Writing Coherent Stories with LLMs — Given a logline, suggest an alternative, original and descriptive title for a known story. That sets the expectation about the input and produces better results using gpt-3.5-turbo. However, I want to stick to the paper as much as possible so the alternative I landed on was to use a dict for the input so the key acts as an implicit label.
- Automatic Story Generation: A Survey of Approaches - ResearchGate — This survey presents an extensive study of research in the area of non-interactive textual story generation, as well as covering resources, corpora, and evaluation methods that have been used in ...
- Guiding and Diversifying LLM-Based Story Generation via Answer Set ... — PDF | Instruction-tuned large language models (LLMs) are capable of generating stories in response to open-ended user requests, but the resulting... | Find, read and cite all the research you need ...
- Wordcraft: Story Writing With Large Language Models - ACM Digital Library — Most of the work described in the previous section has relied on neural language models for generation. Neural language models, such as GPT-2 [] or GPT-Neo [], are neural networks that are trained only to predict the next word in a sequence given the previous words (aka a prompt).We use "large language model," or LLM, to refer to the recent generation of neural language models that have ...
6.2 Open-Source Tools and Libraries
- How to Build a RAG System with Open Source LLMs? — 1.4. Overview of Open Source LLMs. Open Source Large Language Models (LLMs) have gained significant traction in recent years, providing developers and researchers with powerful tools for natural language processing (NLP) tasks. These models are designed to understand and generate human-like text, making them invaluable for various applications.
- A developer's guide to open source LLMs and generative AI — Open source vs. closed source LLMs. By now, most of us are familiar with LLMs: neural network-based language models trained on vast quantities of data to mimic human behavior by performing various downstream tasks, like question answering, translation, and summarization.LLMs have disrupted the world with the introduction of tools like ChatGPT and GitHub Copilot.
- Story Continuation Assistant | AI-powered tool to continue and enhance ... — Provides expert writing advice, assistance, and creative suggestions for ways to continue a story or writing piece. HyperWrite's Story Continuation Assistant is a unique tool that offers expert writing advice and creative suggestions for continuing your stories. Powered by advanced AI models, this tool analyzes your story's style, tone, and plot progression to generate creative and engaging ...
- Exploring Automated Story Generation and Controllable ... - HackerNoon — In "DialogueScript: Using Dialogue Agents to Produce a Script", Schmidtová et al. [96] use different LLMs for each character. Si et al. [102] model multi-user dialogue and character relationships for story continuation. Schmitt and Buschek [97] use question-based chatbot interactions to assist with character creation.
- Open-Source Text Generation & LLM Ecosystem at Hugging Face — Llama is one of the first open-source LLMs to have outperformed/matched closed-source ones. A research group led by Together has created a reproduction of Llama's dataset, called Red Pajama, and trained LLMs and instruction fine-tuned models on it. ... Tools in the Hugging Face Ecosystem for LLM Serving Text Generation Inference
- Guiding and Diversifying LLM-Based Story Generation — 2.1 Plan: Outline Pre-Generation with ASP Figure 1: To generate diverse story outlines via ASP, we first specify a set of high-level narrative functions or storytelling goals that each scene in a story might fulfill. We then define a set of constraints that restrict how these narrative functions are allowed to be sequenced and combined. Finally, using an answer set solver, we solve for the set ...
- Wordcraft: Story Writing With Large Language Models - ACM Digital Library — Participants used open-ended conversation with the chatbot as a method for auditioning story ideas, getting suggestions for story details such as names of characters and locations, and as a more targeted search engine (Section 5.4.2). This is yet another example of the LLM playing a more fluid role for which traditional evaluative metrics such ...
- Story Continuation - Papers With Code — The task involves providing an initial scene that can be obtained in real world use cases. By including this scene, a model can then copy and adapt elements from it as it generates subsequent images. Source: StoryDALL-E: Adapting Pretrained Text-to-Image Transformers for Story Continuation
- Guiding and Diversifying LLM-Based Story Generation via ... - ResearchGate — PDF | Instruction-tuned large language models (LLMs) are capable of generating stories in response to open-ended user requests, but the resulting... | Find, read and cite all the research you need ...
- Twine / An open-source tool for telling interactive, nonlinear stories — Visit the web site of the story format you use to find out how you can support continued development. Non-financial support is important, too. You can help someone out with a question, contribute to the Cookbook , create tutorials of your own, help fix bugs in Twine or its story formats, or translate Twine's user interface to another language.
6.3 Recommended Books and Articles
- Story Continuation AI | Continue your story with AI | HyperWrite AI ... — AI tool to continue a user-provided story, maintaining the style, characters, and plot direction. HyperWrite's Story Continuation AI is an advanced tool that continues your story, maintaining the style, characters, and plot direction. Using cutting-edge AI models, this tool helps you to continue your narrative in a consistent and engaging manner, making it a valuable assistant for writers and ...
- Story Continuation Assistant - HyperWrite | AI — Provides expert writing advice, assistance, and creative suggestions for ways to continue a story or writing piece. HyperWrite's Story Continuation Assistant is a unique tool that offers expert writing advice and creative suggestions for continuing your stories. Powered by advanced AI models, this tool analyzes your story's style, tone, and plot progression to generate creative and engaging ...
- Exploring Automated Story Generation and Controllable ... - HackerNoon — In "DialogueScript: Using Dialogue Agents to Produce a Script", Schmidtová et al. [96] use different LLMs for each character. Si et al. [102] model multi-user dialogue and character relationships for story continuation. Schmitt and Buschek [97] use question-based chatbot interactions to assist with character creation.
- AI Story Generator (free, unlimited, no sign-up) - Perchance — Completely free & unlimited AI story generator/writer based on a prompt. No sign-up or login. Generate LONG stories, paragraph-by-paragraph, optionally guiding the AI on what happens next. Fast generation and there are no daily usage restrictions - unlimited and 100% free storytelling AI, no account needed. You can prompt the AI to create horror stories (including creepy/creepypasta and ...
- Generating Narratives: Writing Coherent Stories with LLMs — Given a logline, suggest an alternative, original and descriptive title for a known story. That sets the expectation about the input and produces better results using gpt-3.5-turbo. However, I want to stick to the paper as much as possible so the alternative I landed on was to use a dict for the input so the key acts as an implicit label.
- Guiding and Diversifying LLM-Based Story Generation — 2.1 Plan: Outline Pre-Generation with ASP Figure 1: To generate diverse story outlines via ASP, we first specify a set of high-level narrative functions or storytelling goals that each scene in a story might fulfill. We then define a set of constraints that restrict how these narrative functions are allowed to be sequenced and combined. Finally, using an answer set solver, we solve for the set ...
- PDF Exploring LLMs and MCTS for Emergent Narrative - ETH Z — In fact, one of the major challenges of working with LLMs is their limitation to work only with unstructured data, while video games rely on structured data. Therefore, in the context of emergent narrative, we propose a hybrid combination method between LLMs and NPC engine systems. Furthermore, we will explore the use of LLMs as a game input method
- Wordcraft: Story Writing With Large Language Models - ACM Digital Library — To study how writers might use LLMs in their work, we conducted a user study in which 25 hobbyist writers were asked to write short stories using Wordcraft. As baselines, we also asked participants to write stories using (1) an AI-powered assistive editor with a single control: continue-my text , and (2) a plain text editor with no extra ...
- PDF Enhancing AI Creativity: A Multi-Agent Approach to Flash Fiction ... — The model integrates two Large Language Models (LLMs): the first, Llama-7B, functions as the story generator which is the primary model undergoing improvement, and the second, gemini-1.5-pro-latest, which acts as the evaluator for prompt fine-tuning. Gemini is tasked with evaluating each draft based
- Source use in the story continuation writing task - ScienceDirect — A similar experiment has been recently conducted in China, a country attaching importance to high-stakes examinations. The story continuation writing task (SCWT) is such a newly developed type of reading-writing integrated task, which is believed to be able to stimulate language learning efficiently (Wang, 2015; Wang & Wang, 2015).







