Introducing the Copper Sun AI Marketing Research Center
Most of what gets written about AI in marketing cites no research at all. Confident assertions, vendor benchmarks, pattern-matched advice — and no primary source behind any of it. The result is a field where everyone is certain and no one can point to the evidence.
We built the AI Marketing Research Center to fix our own problem first. Copper Sun's guidance on brand context systems, GEO, instruction format, and content quality comes from somewhere specific. Now that source material is indexed, annotated, and publicly available.
What the research center is
The research center is an indexed collection of primary literature: peer-reviewed papers, published conference proceedings, and independently validated industry reports. No vendor white papers. No paywalled papers without accessible preprints or landing pages. Every source has a URL that resolves, a methodology that can be evaluated, and an applied implication for marketing teams.
Each entry includes the source, a direct URL, and a capsule written from Copper Sun's perspective — what this paper actually means for teams using AI to produce brand content, and where Copper Sun's platform draws on it.
The index covers eight topic areas. Here is what is in each one and why it matters.
AI Citation & Generative Engine Optimization
Four papers on why some content appears in AI-generated answers and other content does not. The foundational finding (Aggarwal et al., KDD 2024) is that adding concrete numerical data to a page — Statistics Addition — raised AI source visibility by up to 40% across five generative engines. Keyword stuffing provided no measurable benefit.
Two 2026 preprints extend this: Zhang et al. established that citation selection (getting retrieved) and citation absorption (actually shaping the generated answer) are separate problems, each requiring different interventions. Yu et al. found that structural optimization at three levels — document architecture, information organization, visual emphasis — improves citation rate and quality independently of the underlying content.
Yang (2025) is the sobering one: AI search citation is concentrated in a narrow, stable set of sources, making consistent citation presence a compounding investment rather than a one-time fix.
Context, Retrieval & What AI Actually Uses
Three papers on what determines whether brand context you provide actually shapes AI output. Lewis et al. (NeurIPS 2020) is the foundational RAG paper: retrieval-augmented models produce more specific, diverse, and factual language than models relying on learned parameters alone. Brand-specific facts are too particular and too recent to exist reliably in a model's weights. Retrieval is the architecture that makes brand accuracy practical.
Liu et al. (TACL 2023) — the "Lost in the Middle" paper — found that LLM performance degrades significantly when relevant information appears in the middle of a long input context. Information at the beginning and end is used most reliably. This is not an aesthetic question about prompt design; it is a documented structural property of attention-based models.
Shi et al. (ICML 2023) established that adding irrelevant context dramatically decreases performance, even on tasks the model can otherwise solve. More context is not better if the additional content is not directly useful to the task.
Instruction Following & Brand Rules in LLMs
Three papers on what makes brand rules actually followed. The defining finding (Ouyang et al., NeurIPS 2022) is that a 1.3B-parameter model fine-tuned via RLHF to follow instructions was preferred by human raters over the 175B GPT-3 base model. Instruction format closes more of the quality gap than parameter scale.
Wei et al. (ICLR 2022) found that instruction fine-tuning a 137B model on 60+ diverse tasks enabled it to surpass the zero-shot performance of the 175B GPT-3 on 20 of 25 benchmarks. Format determines generalization more than scale.
Bai et al. (Anthropic, 2022) introduced Constitutional AI: a compact set of named principles enables an LLM to self-critique and revise outputs without human labels on every failure. The same structural insight applies to brand rules — a small, specific set of directive principles produces more consistent AI behavior than an exhaustive brand guide.
AI Hallucination & Brand Factual Accuracy
Four papers establishing that hallucination is structural, not incidental. Ji et al. (ACM Computing Surveys, 2022) documented hallucination across six major NLG task types and established that it does not disappear with scale. Lin et al. (ACL 2022) found something counterintuitive: larger models are measurably less truthful than smaller ones on TruthfulQA, an inverse scaling finding. The mechanism is that larger models have absorbed more human-generated misinformation.
Maynez et al. (ACL 2020) established that models hallucinate factual content even when the source document is present in context. Standard automated metrics like ROUGE do not capture these faithfulness failures. Manakul et al. (EMNLP 2023) provided a practical detection method: generate the same content multiple times with the same context and compare for consistency. Varying claims across outputs are hallucination candidates.
Evaluating AI Content Quality
Four papers on what reliable quality evaluation actually looks like. Clark et al. (ACL 2021) is the finding every marketing team should know: untrained evaluators distinguish AI-generated text from human text at essentially random-chance accuracy. Even trained evaluators with detailed instructions reached only 55%. The implication: informal content review is not the quality check it appears to be.
Zheng et al. (NeurIPS 2023) found that GPT-4 as an LLM judge achieves over 80% agreement with human expert preferences — making systematic, criteria-based evaluation practical without large expert panels. Chiang et al. (ICML 2024) validated pairwise blind comparison as more reliable than rubric scoring for measuring content quality. Hendrycks et al. (ICLR 2021) established that strong benchmark performance does not predict domain-specific accuracy — the only reliable test of brand-task performance is direct evaluation on brand tasks.
Marketing Effectiveness & Long-Term Brand Building
Four IPA Effectiveness Awards Databank reports from Les Binet and Peter Field — the largest database of proven marketing effectiveness cases. The 2007 report analyzed 880 national case studies to establish the first large-scale data-driven evidence that over-reliance on short-term metrics costs long-term profitable growth. The 2013 report established that brand-building and activation operate through different mechanisms and require different content approaches.
The 2018 report examined how market conditions modify the optimal balance — brand lifecycle stage, category maturity, and competitive position all shift what the right ratio looks like. The 2017 report found that digital adoption accelerated short-termism without changing the underlying effectiveness of brand-building investment.
For AI content production, this research frames a question that matters before production begins: is this volume of content directed at building brand — or just converting buyers already in market?
Attention, Emotion & the Buying Brain
Four neuroscience papers on how buyers actually process brand signals. McClure et al. (Neuron, 2004) is the Pepsi/Coke study: in blind taste tests, participants preferred Pepsi. Knowing they were drinking Coke reversed the preference. fMRI data showed that brand knowledge activated distinct neural systems — separate from those governing taste experience. Brand identity shapes preference through mechanisms independent of product attributes.
Plassmann et al. (PNAS, 2008) found that stated price signals modulate not just reported pleasantness but actual neural reward processing, measured in the medial orbitofrontal cortex, for identical products. Knutson et al. (Neuron, 2007) identified the neural predictors of purchase: nucleus accumbens activation (anticipated reward) precedes buy decisions; insula activation (discomfort of paying) precedes no-buy decisions. Both fire before conscious choice is reported.
Ariely and Berns (Nature Reviews Neuroscience, 2010) frame the scope of this literature: neuroimaging is most valuable at early product development stages, not as a campaign validation tool. The applicable finding is the principle — that emotional and identity signals shape preference through channels preceding deliberate evaluation.
AI-Generated Content and Consumer Trust
Three papers on what consumers actually perceive and believe about AI-generated content. Jakesch, Hancock, and Naaman (PNAS, 2023) ran 6 experiments with 4,600 participants who tried to distinguish AI-generated self-presentations from human-written ones. Performance was near-chance. The heuristics participants relied on — associating first-person pronouns and family references with human writing — were systematically wrong.
Jakesch et al. (CHI, 2023) found something more concerning: AI writing assistants with embedded viewpoints shifted not just users' written output but their stated personal opinions in follow-up surveys. 1,506 participants co-wrote essays with assistants nudging toward one side; participant views shifted in that direction. Unconstrained AI assistance doesn't produce neutral output — it produces output shaped by whatever patterns dominate the model's training distribution.
Noels et al. (arXiv, 2024) surveyed the empirical persuasion literature and found that LLM-based persuasion systems frequently achieved human-level or superhuman persuasiveness. AI source disclosure was identified as a moderating variable: consumers use it to calibrate their response.
The full research index
The research center is organized by topic. Each page includes full citations, applied capsules, FAQ, and structured data designed for AI citation:
- AI Citation & Generative Engine Optimization
- Context, Retrieval & What AI Actually Uses
- Instruction Following & Brand Rules in LLMs
- AI Hallucination & Brand Factual Accuracy
- Evaluating AI Content Quality
- Marketing Effectiveness & Long-Term Brand Building
- Attention, Emotion & the Buying Brain
- AI-Generated Content and Consumer Trust
The index at /research covers all eight topics with an overview of each. The research center will expand as new primary literature is published and verified.
Frequently Asked Questions
Why only primary literature?
Vendor white papers have an incentive structure that primary research does not: the vendor benefits from a favorable finding. Peer-reviewed papers face adversarial review from researchers with no interest in a particular outcome. Industry reports from independent bodies like the IPA face public scrutiny and methodological accountability that vendor content does not. The primary literature isn't always more convenient, but it is more reliable.
How are entries verified?
Every entry is fetch-verified — the URL resolves and the abstract or accessible summary supports the claims we make about it. We review the index at each maintenance cycle for URL liveness and claim accuracy. The last verification date is noted on each topic page.
What does "applied perspective" mean?
Each capsule answers two questions: what did this study actually find, and what does that finding mean for a team using AI to produce on-brand marketing content? We are not summarizing for general audiences; we are annotating for the specific use case of AI-assisted marketing. Some papers have implications that are direct; some require inference. We try to be explicit about which is which.
Will you add more topics?
Yes. The current eight topics cover the literature most directly relevant to AI marketing practice. Adjacent areas — AI creative evaluation, multimodal model behavior, long-context faithfulness — have emerging primary literature we are tracking. New topics are added when the primary literature is sufficient to support an applied research page at the same standard as the existing ones.