<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://diogoribeiro7.github.io/analytics-blog-jekyll/feed.xml" rel="self" type="application/atom+xml" /><link href="https://diogoribeiro7.github.io/analytics-blog-jekyll/" rel="alternate" type="text/html" /><updated>2026-09-27T23:42:02+01:00</updated><id>https://diogoribeiro7.github.io/analytics-blog-jekyll/feed.xml</id><title type="html">DataLog | Data Science &amp;amp; Research Theme</title><subtitle>DataLog is a clean and academic-inspired Jekyll theme crafted for data scientists, researchers, and technical writers who want to share reproducible analyses, research papers, tutorials, datasets, and project portfolios.</subtitle><author><name>Diogo Ribeiro</name><email>dfr@esmad.ipp.pt</email></author><entry><title type="html">Minimal Mistakes Front Matter, Rendered by DataLog</title><link href="https://diogoribeiro7.github.io/analytics-blog-jekyll/tutorials/2026/05/05/minimal-mistakes-front-matter-compatibility/" rel="alternate" type="text/html" title="Minimal Mistakes Front Matter, Rendered by DataLog" /><published>2026-05-05T00:00:00+01:00</published><updated>2026-05-05T00:00:00+01:00</updated><id>https://diogoribeiro7.github.io/analytics-blog-jekyll/tutorials/2026/05/05/minimal-mistakes-front-matter-compatibility</id><content type="html" xml:base="https://diogoribeiro7.github.io/analytics-blog-jekyll/tutorials/2026/05/05/minimal-mistakes-front-matter-compatibility/"><![CDATA[<p>This post's front matter is written the way the <a href="https://mmistakes.github.io/minimal-mistakes/">Minimal Mistakes</a> theme expects it, and none of it has been renamed for DataLog. The hero image above, the teaser on the home page cards, the title in the browser tab, the meta description, the wide body class and the editorial note below all come from Minimal Mistakes fields.</p>

<h2 id="what-was-mapped">What was mapped</h2>

<div class="content-table" role="region" tabindex="0" aria-label="Table 1"><table>
  <thead>
    <tr>
      <th>Minimal Mistakes field</th>
      <th>DataLog field</th>
      <th>Where it shows</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">header.overlay_image</code></td>
      <td><code class="language-plaintext highlighter-rouge">post_hero.image</code> with <code class="language-plaintext highlighter-rouge">post_hero.overlay: true</code></td>
      <td>The title sits over the image at the top of this page</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">header.overlay_filter</code></td>
      <td><code class="language-plaintext highlighter-rouge">post_hero.overlay_filter</code></td>
      <td>The darkening over the hero image</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">header.teaser</code></td>
      <td><code class="language-plaintext highlighter-rouge">teaser</code></td>
      <td>Thumbnail on the home page and related-post cards</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">header.og_image</code>, <code class="language-plaintext highlighter-rouge">header.twitter_image</code></td>
      <td><code class="language-plaintext highlighter-rouge">og_image</code>, <code class="language-plaintext highlighter-rouge">twitter_image</code></td>
      <td>Social sharing previews</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">seo_title</code></td>
      <td><code class="language-plaintext highlighter-rouge">&lt;title&gt;</code> and social titles</td>
      <td>The browser tab reads the SEO title, the heading reads the real one</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">seo_description</code></td>
      <td><code class="language-plaintext highlighter-rouge">description</code></td>
      <td>The meta description</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">classes: wide</code></td>
      <td>body class <code class="language-plaintext highlighter-rouge">wide</code></td>
      <td>Paragraphs use the full content width</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">subtitle</code></td>
      <td>subtitle under the title</td>
      <td>The line under the heading</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">why_this_exists</code>, <code class="language-plaintext highlighter-rouge">evidence</code>, <code class="language-plaintext highlighter-rouge">methodology</code>, <code class="language-plaintext highlighter-rouge">reviewed_at</code></td>
      <td>provenance note</td>
      <td>The editorial note above the article body</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">redirect_from</code></td>
      <td>redirect page</td>
      <td><code class="language-plaintext highlighter-rouge">/legacy/minimal-mistakes-post/</code> redirects here</td>
    </tr>
  </tbody>
</table></div>

<h2 id="what-is-ignored">What is ignored</h2>

<p><code class="language-plaintext highlighter-rouge">author_profile</code> controls the Minimal Mistakes sidebar, which DataLog does not have; DataLog shows the author card at the end of the post instead. <code class="language-plaintext highlighter-rouge">seo_type: article</code> is redundant, because DataLog already marks posts as articles for social cards and structured data.</p>

<h2 id="where-the-rules-live">Where the rules live</h2>

<p>The mapping is a single build hook in <code class="language-plaintext highlighter-rouge">_plugins/front_matter_compat.rb</code>. It only fills DataLog fields that are absent, so a post that sets <code class="language-plaintext highlighter-rouge">image</code> or <code class="language-plaintext highlighter-rouge">description</code> explicitly keeps those values. The full field table and the URL-preservation settings are in the <a href="https://github.com/DiogoRibeiro7/analytics-blog-jekyll/blob/develop/docs/migrating-from-minimal-mistakes.md">migration guide</a>.</p>]]></content><author><name>Diogo Ribeiro</name></author><category term="tutorials" /><category term="migration" /><category term="minimal-mistakes" /><category term="front-matter" /><summary type="html"><![CDATA[How DataLog reads Minimal Mistakes header images, teasers, SEO titles and descriptions, wide layouts and redirects without editing every post.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://diogoribeiro7.github.io/analytics-blog-jekyll/assets/img/social-card.png" /><media:content medium="image" url="https://diogoribeiro7.github.io/analytics-blog-jekyll/assets/img/social-card.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Dataset Announcement – Retail Demand Benchmark</title><link href="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/11/dataset-announcement-schema/" rel="alternate" type="text/html" title="Dataset Announcement – Retail Demand Benchmark" /><published>2024-04-11T00:00:00+01:00</published><updated>2024-04-11T00:00:00+01:00</updated><id>https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/11/dataset-announcement-schema</id><content type="html" xml:base="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/11/dataset-announcement-schema/"><![CDATA[<p>Today we are releasing the <strong>Retail Demand Benchmark</strong> dataset to help analytics teams evaluate forecasting models under realistic constraints. The DataLog theme lets you pair documentation, schema tables, and governance checklists in one place.</p>

<h2 id="download-options">Download options</h2>

<ul>
  <li><code class="language-plaintext highlighter-rouge">retail-demand-benchmark.csv</code> via the dataset page</li>
  <li>Parquet export via <code class="language-plaintext highlighter-rouge">_datasets/retail-demand-benchmark.md</code></li>
  <li>Companion notebook: <code class="language-plaintext highlighter-rouge">_notebooks/batch-anomaly-detection.ipynb</code></li>
</ul>

<h2 id="schema">Schema</h2>

<div class="content-table" role="region" tabindex="0" aria-label="Table 1"><table>
  <thead>
    <tr>
      <th>Column</th>
      <th>Type</th>
      <th>Description</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">date</code></td>
      <td>Date</td>
      <td>Week-ending timestamp</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">store_id</code></td>
      <td>String</td>
      <td>Unique store identifier</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">category</code></td>
      <td>String</td>
      <td>Product category slug</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">units_sold</code></td>
      <td>Integer</td>
      <td>Units sold in the week</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">price</code></td>
      <td>Decimal</td>
      <td>Average selling price</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">promo_flag</code></td>
      <td>Boolean</td>
      <td>Promotion indicator</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">stockout_minutes</code></td>
      <td>Integer</td>
      <td>Minutes unavailable due to stockouts</td>
    </tr>
  </tbody>
</table></div>

<h2 id="usage-guidelines">Usage guidelines</h2>

<div class="language-yaml highlighter-rouge"><div class="highlight"><pre class="highlight" tabindex="0"><code><table class="rouge-table"><tbody><tr><td class="rouge-gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
</pre></td><td class="rouge-code"><pre><span class="na">data_quality</span><span class="pi">:</span>
  <span class="na">freshness</span><span class="pi">:</span> <span class="s">P1D</span>
  <span class="na">source_system</span><span class="pi">:</span> <span class="s">retail_warehouse</span>
  <span class="na">owners</span><span class="pi">:</span>
    <span class="pi">-</span> <span class="na">name</span><span class="pi">:</span> <span class="s">Diogo Ribeiro</span>
      <span class="na">email</span><span class="pi">:</span> <span class="s">diogo.debastos.ribeiro@gmail.com</span>
    <span class="pi">-</span> <span class="na">name</span><span class="pi">:</span> <span class="s">Analytics Platform Team</span>
      <span class="na">email</span><span class="pi">:</span> <span class="s">platform@example.com</span>
  <span class="na">sla</span><span class="pi">:</span>
    <span class="na">anomalies</span><span class="pi">:</span> <span class="s">notify-analytics-oncall@example.com</span>
    <span class="na">completeness</span><span class="pi">:</span> <span class="pi">&gt;</span><span class="err">=</span> <span class="err">0.98</span>
</pre></td></tr></tbody></table></code></pre></div></div>

<blockquote>
  <p><strong>Governance tip:</strong> Store ownership metadata in <code class="language-plaintext highlighter-rouge">_data/catalog.yml</code> so the theme can surface it across dataset and notebook pages.</p>
</blockquote>

<h2 id="next-steps">Next steps</h2>

<ol>
  <li>Publish baseline forecasts and accuracy benchmarks in the notebooks collection.</li>
  <li>Embed interactive Plotly charts showing store-level variability.</li>
  <li>Solicit community contributions via GitHub Discussions and credit accepted pull requests in the showcase gallery.</li>
</ol>

<p>With schema documentation and ownership metadata in place, teams can trust the dataset before wiring it into production pipelines.</p>]]></content><author><name>Diogo Ribeiro</name></author><category term="dataset" /><category term="announcement" /><category term="documentation" /><summary type="html"><![CDATA[Today we are releasing the Retail Demand Benchmark dataset to help analytics teams evaluate forecasting models under realistic constraints. The DataLog theme lets you pair documentation, schema tables, and governance checklists in one place.]]></summary></entry><entry><title type="html">Disqus CSP Verification</title><link href="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/10/disqus-csp-verification/" rel="alternate" type="text/html" title="Disqus CSP Verification" /><published>2024-04-10T00:00:00+01:00</published><updated>2024-04-10T00:00:00+01:00</updated><id>https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/10/disqus-csp-verification</id><content type="html" xml:base="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/10/disqus-csp-verification/"><![CDATA[<p>This hidden post exists purely for automated tests that confirm inline scripts emitted by the
Disqus embed receive the correct Content Security Policy nonce. It is excluded from the primary
navigation but published so that the Jekyll site generator renders the markup used by the tests.</p>]]></content><author><name>Diogo Ribeiro</name></author><category term="testing" /><summary type="html"><![CDATA[This hidden post exists purely for automated tests that confirm inline scripts emitted by the Disqus embed receive the correct Content Security Policy nonce. It is excluded from the primary navigation but published so that the Jekyll site generator renders the markup used by the tests.]]></summary></entry><entry><title type="html">Portfolio Case Study with GitHub Insights</title><link href="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/10/portfolio-case-study-github-stats/" rel="alternate" type="text/html" title="Portfolio Case Study with GitHub Insights" /><published>2024-04-10T00:00:00+01:00</published><updated>2024-04-10T00:00:00+01:00</updated><id>https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/10/portfolio-case-study-github-stats</id><content type="html" xml:base="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/10/portfolio-case-study-github-stats/"><![CDATA[<p>Case studies shine when readers can verify the underlying code. This example highlights a recommendation engine project, pulling GitHub stats directly into the narrative and linking supporting resources.</p>

<h2 id="project-overview">Project overview</h2>

<ul>
  <li><strong>Client:</strong> Streaming media startup scaling personalization</li>
  <li><strong>Role:</strong> Lead machine learning engineer</li>
  <li><strong>Timeline:</strong> Q1 2024 (8 weeks)</li>
</ul>

<h2 id="repository-snapshot">Repository snapshot</h2>

<p><img src="https://img.shields.io/github/stars/DiogoRibeiro7/datalog-starter?style=social" alt="GitHub stars">
<img src="https://img.shields.io/github/forks/DiogoRibeiro7/datalog-starter?style=social" alt="GitHub forks">
<img src="https://img.shields.io/github/last-commit/DiogoRibeiro7/datalog-starter" alt="GitHub last commit"></p>

<blockquote>
  <p><strong>Why it matters:</strong> Badges refresh automatically so stakeholders always see the latest adoption metrics.</p>
</blockquote>

<h2 id="architecture">Architecture</h2>

<pre><code class="language-mermaid">graph TD
  A[User events] --&gt; B[Feature store]
  B --&gt; C[Candidate generation]
  C --&gt; D[Ranking model]
  D --&gt; E[Realtime API]
  E --&gt; F[Personalized UI]
</code></pre>

<h2 id="outcomes">Outcomes</h2>

<div class="content-table" role="region" tabindex="0" aria-label="Table 1"><table>
  <thead>
    <tr>
      <th>KPI</th>
      <th>Baseline</th>
      <th>Post-launch</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Recommendation CTR</td>
      <td>4.8%</td>
      <td><strong>6.0%</strong></td>
    </tr>
    <tr>
      <td>Hours streamed / user</td>
      <td>5.1</td>
      <td><strong>6.3</strong></td>
    </tr>
    <tr>
      <td>Support tickets / week</td>
      <td>18</td>
      <td><strong>11</strong></td>
    </tr>
  </tbody>
</table></div>

<h2 id="supporting-assets">Supporting assets</h2>

<ul>
  <li><code class="language-plaintext highlighter-rouge">_notebooks/recsys-evaluation.ipynb</code> – Offline evaluation and fairness analysis</li>
  <li><code class="language-plaintext highlighter-rouge">_datasets/recsys-segment-metrics.csv</code> – Segment-level KPI tracking</li>
  <li><code class="language-plaintext highlighter-rouge">_portfolio/market-segmentation-case-study.md</code> – Extended client background</li>
</ul>

<p>Wrap up by inviting readers to clone the repo, run the notebooks, and explore the dashboards embedded elsewhere on the site.</p>]]></content><author><name>Diogo Ribeiro</name></author><category term="portfolio" /><category term="case-study" /><category term="github" /><summary type="html"><![CDATA[Case studies shine when readers can verify the underlying code. This example highlights a recommendation engine project, pulling GitHub stats directly into the narrative and linking supporting resources.]]></summary></entry><entry><title type="html">Experimental Design Blueprint with Power Analysis</title><link href="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/09/experimental-design-statistical-tests/" rel="alternate" type="text/html" title="Experimental Design Blueprint with Power Analysis" /><published>2024-04-09T00:00:00+01:00</published><updated>2024-04-09T00:00:00+01:00</updated><id>https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/09/experimental-design-statistical-tests</id><content type="html" xml:base="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/09/experimental-design-statistical-tests/"><![CDATA[<p>Thoughtful experiment logs help product teams align on hypotheses before shipping features. This blueprint captures the essentials—design tables, power analysis code, and interpretation guidelines—all rendered cleanly by the <strong>DataLog</strong> theme.</p>

<h2 id="hypotheses">Hypotheses</h2>

<div class="content-table" role="region" tabindex="0" aria-label="Table 1"><table>
  <thead>
    <tr>
      <th>Hypothesis</th>
      <th>Description</th>
      <th>Metric</th>
      <th>Direction</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>H1</td>
      <td>New onboarding improves activation</td>
      <td>Activation rate</td>
      <td>Increase</td>
    </tr>
    <tr>
      <td>H2</td>
      <td>Tooltips reduce setup time</td>
      <td>Median time-to-value</td>
      <td>Decrease</td>
    </tr>
  </tbody>
</table></div>

<h2 id="sample-size-planning">Sample-size planning</h2>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight" tabindex="0"><code><table class="rouge-table"><tbody><tr><td class="rouge-gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
12
13
</pre></td><td class="rouge-code"><pre><span class="kn">from</span> <span class="n">statsmodels.stats.power</span> <span class="kn">import</span> <span class="n">NormalIndPower</span>

<span class="n">effect</span> <span class="o">=</span> <span class="mf">0.04</span>  <span class="c1"># minimum detectable effect (absolute)
</span><span class="n">alpha</span> <span class="o">=</span> <span class="mf">0.05</span>
<span class="n">power</span> <span class="o">=</span> <span class="mf">0.8</span>
<span class="n">baseline</span> <span class="o">=</span> <span class="mf">0.32</span>

<span class="n">analysis</span> <span class="o">=</span> <span class="nc">NormalIndPower</span><span class="p">()</span>
<span class="n">n_per_group</span> <span class="o">=</span> <span class="n">analysis</span><span class="p">.</span><span class="nf">solve_power</span><span class="p">(</span><span class="n">effect_size</span><span class="o">=</span><span class="n">effect</span> <span class="o">/</span> <span class="p">(</span><span class="n">baseline</span> <span class="o">*</span> <span class="p">(</span><span class="mi">1</span> <span class="o">-</span> <span class="n">baseline</span><span class="p">))</span> <span class="o">**</span> <span class="mf">0.5</span><span class="p">,</span>
                                  <span class="n">power</span><span class="o">=</span><span class="n">power</span><span class="p">,</span>
                                  <span class="n">alpha</span><span class="o">=</span><span class="n">alpha</span><span class="p">,</span>
                                  <span class="n">ratio</span><span class="o">=</span><span class="mf">1.0</span><span class="p">)</span>
<span class="nf">print</span><span class="p">(</span><span class="nf">round</span><span class="p">(</span><span class="n">n_per_group</span><span class="p">))</span>
</pre></td></tr></tbody></table></code></pre></div></div>

<blockquote>
  <p><strong>Reminder:</strong> Adjust for multiple comparisons if you expect to peek at intermediate checkpoints.</p>
</blockquote>

<h2 id="test-plan">Test plan</h2>

<div class="content-table" role="region" tabindex="0" aria-label="Table 2"><table>
  <thead>
    <tr>
      <th>Metric</th>
      <th>Test</th>
      <th>Rationale</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Activation rate</td>
      <td>Two-proportion z-test</td>
      <td>Large samples, binary outcome</td>
    </tr>
    <tr>
      <td>Time-to-value</td>
      <td>Mann–Whitney U</td>
      <td>Non-parametric, skewed distribution</td>
    </tr>
    <tr>
      <td>Retention (D28)</td>
      <td>Kaplan–Meier log-rank</td>
      <td>Survival analysis</td>
    </tr>
  </tbody>
</table></div>

<h2 id="decision-framework">Decision framework</h2>

<ol>
  <li>Pre-register hypotheses and guardrails in <code class="language-plaintext highlighter-rouge">_datasets/experiment-hypotheses.csv</code>.</li>
  <li>Automate metric extraction via notebooks stored in <code class="language-plaintext highlighter-rouge">_notebooks/</code>.</li>
  <li>Attach Tableau or Looker dashboards with the <code class="language-plaintext highlighter-rouge">viz-block</code> include for executive readouts.</li>
</ol>

<p>By combining Markdown tables, statistical code, and callouts, the post becomes a reusable experimentation template for every squad.</p>]]></content><author><name>Diogo Ribeiro</name></author><category term="experimentation" /><category term="statistics" /><category term="design" /><summary type="html"><![CDATA[Thoughtful experiment logs help product teams align on hypotheses before shipping features. This blueprint captures the essentials—design tables, power analysis code, and interpretation guidelines—all rendered cleanly by the DataLog theme.]]></summary></entry><entry><title type="html">Mathematical Proof Template with Numbered Equations</title><link href="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/08/mathematical-proof-numbered-equations/" rel="alternate" type="text/html" title="Mathematical Proof Template with Numbered Equations" /><published>2024-04-08T00:00:00+01:00</published><updated>2024-04-08T00:00:00+01:00</updated><id>https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/08/mathematical-proof-numbered-equations</id><content type="html" xml:base="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/08/mathematical-proof-numbered-equations/"><![CDATA[<p>Document rigorous mathematics directly inside your documentation hub. This template illustrates how theorem statements, lemmas, and numbered equations render cleanly on the <strong>DataLog</strong> theme.</p>

<h2 id="theorem">Theorem</h2>

<blockquote>
  <p><strong>Theorem 1.</strong> Let <span class="math-expression-inline math-expression--source" role="math" data-math-alt="f : [a, b] R" data-math-source="f : [a, b] \to \mathbb{R}" aria-label="f : [a, b] R" tabindex="0">$f : [a, b] \to \mathbb{R}$</span> be twice continuously differentiable with <span class="math-expression-inline math-expression--source" role="math" data-math-alt="f(a) = f(b) = 0" data-math-source="f(a) = f(b) = 0" aria-label="f(a) = f(b) = 0" tabindex="0">$f(a) = f(b) = 0$</span>. Then there exists <span class="math-expression-inline math-expression--source" role="math" data-math-alt="c (a, b)" data-math-source="c \in (a, b)" aria-label="c (a, b)" tabindex="0">$c \in (a, b)$</span> such that <span class="math-expression-inline math-expression--source" role="math" data-math-alt="f&#39;&#39;(c) + ^2 over (b-a)^2 f(c) = 0" data-math-source="f&#39;&#39;(c) + \frac{\pi^2}{(b-a)^2} f(c) = 0" aria-label="f&#39;&#39;(c) + ^2 over (b-a)^2 f(c) = 0" tabindex="0">$f&#39;&#39;(c) + \frac{\pi^2}{(b-a)^2} f(c) = 0$</span>.</p>
</blockquote>

<h2 id="proof">Proof</h2>

<p>We adapt the standard Wirtinger inequality. Consider the sine basis function <span class="math-expression-inline math-expression--source" role="math" data-math-alt="g(x) = ((x-a) over b-a )" data-math-source="g(x) = \sin\left(\frac{\pi(x-a)}{b-a}\right)" aria-label="g(x) = ((x-a) over b-a )" tabindex="0">$g(x) = \sin\left(\frac{\pi(x-a)}{b-a}\right)$</span> and define</p>

<div class="math-expression math-expression--source" role="math" data-math-alt="inner product equals the integral of f times g over the interval from a to b" data-math-source="\label{eq:inner-product}
% alt: inner product equals the integral of f times g over the interval from a to b
\langle f, g \rangle = \int_a^b f(x) g(x) \, \mathrm{d}x." aria-label="inner product equals the integral of f times g over the interval from a to b" tabindex="0">\begin{equation}\label{eq:inner-product}
% alt: inner product equals the integral of f times g over the interval from a to b
\langle f, g \rangle = \int_a^b f(x) g(x) \, \mathrm{d}x.
\end{equation}</div>

<p>Integration by parts shows that</p>

<div class="math-expression math-expression--source" role="math" data-math-alt="eq:ibp integral from a to b f&#39;(x) g&#39;(x) \, dx = -integral from a to b f(x) g&#39;&#39;(x) \, dx = ^2 over (b-a)^2 f, g ." data-math-source="\label{eq:ibp}
\int_a^b f&#39;(x) g&#39;(x) \, \mathrm{d}x = -\int_a^b f(x) g&#39;&#39;(x) \, \mathrm{d}x = \frac{\pi^2}{(b-a)^2} \langle f, g \rangle." aria-label="eq:ibp integral from a to b f&#39;(x) g&#39;(x) \, dx = -integral from a to b f(x) g&#39;&#39;(x) \, dx = ^2 over (b-a)^2 f, g ." tabindex="0">\begin{equation}\label{eq:ibp}
\int_a^b f&#39;(x) g&#39;(x) \, \mathrm{d}x = -\int_a^b f(x) g&#39;&#39;(x) \, \mathrm{d}x = \frac{\pi^2}{(b-a)^2} \langle f, g \rangle.
\end{equation}</div>

<p>Combining Equations \eqref{eq:inner-product} and \eqref{eq:ibp} yields the desired critical point when <span class="math-expression-inline math-expression--source" role="math" data-math-alt="f" data-math-source="f" aria-label="f" tabindex="0">$f$</span> is not identically zero. <span class="math-expression-inline math-expression--source" role="math" data-math-alt="Mathematical expression" data-math-source="\square" aria-label="Mathematical expression" tabindex="0">$\square$</span></p>

<h2 id="notes-for-authors">Notes for authors</h2>

<ul>
  <li>MathJax automatically numbers <code class="language-plaintext highlighter-rouge">equation</code> environments so you can reference them with <code class="language-plaintext highlighter-rouge">\eqref{}</code>.</li>
  <li>Inline math, such as <span class="math-expression-inline math-expression--source" role="math" data-math-alt="integral from a to b f(x)\,dx" data-math-source="\int_a^b f(x)\,\mathrm{d}x" aria-label="integral from a to b f(x)\,dx" tabindex="0">$\int_a^b f(x)\,\mathrm{d}x$</span>, remains crisp across light and dark modes.</li>
  <li>Use definition, lemma, and corollary blocks as needed—Markdown blockquotes keep the typography consistent.</li>
</ul>

<p>Add proof sketches, exercises, or downloadable solution PDFs to round out your mathematical articles.</p>]]></content><author><name>Diogo Ribeiro</name></author><category term="mathematics" /><category term="proof" /><category term="latex" /><summary type="html"><![CDATA[Document rigorous mathematics directly inside your documentation hub. This template illustrates how theorem statements, lemmas, and numbered equations render cleanly on the DataLog theme.]]></summary></entry><entry><title type="html">Plotly Visualization Showcase for Executive Dashboards</title><link href="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/07/data-visualization-plotly-showcase/" rel="alternate" type="text/html" title="Plotly Visualization Showcase for Executive Dashboards" /><published>2024-04-07T00:00:00+01:00</published><updated>2024-04-07T00:00:00+01:00</updated><id>https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/07/data-visualization-plotly-showcase</id><content type="html" xml:base="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/07/data-visualization-plotly-showcase/"><![CDATA[<p>Interactive dashboards help executives interrogate results without leaving the page. The <strong>DataLog</strong> theme ships with a <code class="language-plaintext highlighter-rouge">viz-block</code> component that wraps Plotly embeds with metadata, status messaging, and export buttons.</p>

<h2 id="revenue-trends">Revenue trends</h2>

<div class="viz-block" data-viz-type="plotly" data-viz-slug="plotly-revenue-trends" data-viz-version="1.0" data-viz-updated="2024-04-07">
  <div class="viz-header">
    <h3 class="viz-title" id="quarterly-revenue-growth">Quarterly revenue growth</h3>
    <p class="viz-meta" data-viz-meta="">ARR segmented by region and channel</p>
    <span class="viz-status" data-viz-status="" aria-live="polite">Loading…</span>
  </div>
  <div class="viz-toolbar" role="group" aria-label="Plotly controls">
    <div data-viz-export=""></div>
    <a class="viz-toolbar__button" href="https://plotly.com/javascript/">Plotly docs</a>
  </div>
  <div class="viz-canvas" data-viz-canvas="">
    <script nonce="SgEFAC/1lmVMu/ixdE1bMw==" type="application/json">
{
  "data": [
    {"type": "bar", "name": "North America", "x": ["Q1", "Q2", "Q3", "Q4"], "y": [4.2, 4.6, 5.1, 5.7]},
    {"type": "bar", "name": "EMEA", "x": ["Q1", "Q2", "Q3", "Q4"], "y": [3.1, 3.4, 3.9, 4.3]},
    {"type": "bar", "name": "APAC", "x": ["Q1", "Q2", "Q3", "Q4"], "y": [2.5, 2.9, 3.3, 3.6]}
  ],
  "layout": {
    "barmode": "group",
    "title": {"text": "Revenue by region (in millions USD)"},
    "yaxis": {"title": "Revenue"},
    "legend": {"orientation": "h", "x": 0.5, "xanchor": "center"},
    "margin": {"t": 64, "r": 32, "b": 64, "l": 56}
  }
}
    </script>
  </div>
</div>

<h2 id="why-it-matters">Why it matters</h2>

<ul>
  <li>Keyboard users gain export controls via the toolbar.</li>
  <li>Metadata automatically exposes the chart title and version in screen reader-friendly text.</li>
  <li>Embedding raw JSON keeps version diffs tight during code review.</li>
</ul>

<p>Pair this Plotly block with D3, Observable, and Shiny embeds to demonstrate the full visualization gallery offered by the theme.</p>]]></content><author><name>Diogo Ribeiro</name></author><category term="visualization" /><category term="plotly" /><category term="dashboard" /><summary type="html"><![CDATA[Interactive dashboards help executives interrogate results without leaving the page. The DataLog theme ships with a viz-block component that wraps Plotly embeds with metadata, status messaging, and export buttons.]]></summary></entry><entry><title type="html">Research Article Template with Citations and BibTeX</title><link href="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/06/research-paper-with-citations/" rel="alternate" type="text/html" title="Research Article Template with Citations and BibTeX" /><published>2024-04-06T00:00:00+01:00</published><updated>2024-04-06T00:00:00+01:00</updated><id>https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/06/research-paper-with-citations</id><content type="html" xml:base="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/06/research-paper-with-citations/"><![CDATA[<p>Publishing reproducible scholarship requires more than compelling charts. This template demonstrates how to structure a research article, cite related work, and provide BibTeX metadata so colleagues can reference your study quickly.</p>

<h2 id="abstract">Abstract</h2>

<p>We evaluate adaptive experimentation for recommendation systems, focusing on policy regret minimization across cold-start cohorts. Empirical results indicate a 12% lift in engagement relative to static baselines while maintaining fairness constraints.</p>

<h2 id="introduction">Introduction</h2>

<p>Personalized experiences need to balance accuracy and fairness. Prior work on contextual bandits<sup id="fnref:1"><a href="#fn:1" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> and constrained optimization<sup id="fnref:2"><a href="#fn:2" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> lays the foundation for our framework.</p>

<h2 id="methodology">Methodology</h2>

<p>We define policy regret as</p>

<div class="math-expression math-expression--source" role="math" data-math-alt="eq:regret R _T = summation from t=1 to T ( r_t(x_t, a_t^ ) - r_t(x_t, a_t) )" data-math-source="\label{eq:regret}
\mathcal{R}_T = \sum_{t=1}^T \bigl( r_t(x_t, a_t^\star) - r_t(x_t, a_t) \bigr)" aria-label="eq:regret R _T = summation from t=1 to T ( r_t(x_t, a_t^ ) - r_t(x_t, a_t) )" tabindex="0">\begin{equation}\label{eq:regret}
\mathcal{R}_T = \sum_{t=1}^T \bigl( r_t(x_t, a_t^\star) - r_t(x_t, a_t) \bigr)
\end{equation}</div>

<p>where <span class="math-expression-inline math-expression--source" role="math" data-math-alt="r_t" data-math-source="r_t" aria-label="r_t" tabindex="0">$r_t$</span> is the reward and <span class="math-expression-inline math-expression--source" role="math" data-math-alt="a_t^" data-math-source="a_t^\star" aria-label="a_t^" tabindex="0">$a_t^\star$</span> is the action chosen by an oracle. Algorithm 1 summarizes the constrained Thompson sampling procedure.</p>

<pre><code class="language-pseudo">Initialize posterior priors for all arms
for each round t = 1..T:
  sample reward estimates from posterior
  project samples to satisfy fairness constraints
  choose arm with highest adjusted draw
  update posterior with observed reward
</code></pre>

<h2 id="results">Results</h2>

<div class="content-table" role="region" tabindex="0" aria-label="Table 1"><table>
  <thead>
    <tr>
      <th>Metric</th>
      <th>Baseline</th>
      <th>Adaptive policy</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Click-through rate</td>
      <td>5.4%</td>
      <td><strong>6.1%</strong></td>
    </tr>
    <tr>
      <td>Retention (28-day)</td>
      <td>42.0%</td>
      <td><strong>45.8%</strong></td>
    </tr>
    <tr>
      <td>Fairness gap (Δ)</td>
      <td>0.17</td>
      <td><strong>0.06</strong></td>
    </tr>
  </tbody>
</table></div>

<h2 id="discussion">Discussion</h2>

<p>Equation \eqref{eq:regret} highlights how regret decomposes into reward differences. Future work will incorporate causal constraints to prevent drift.</p>

<h2 id="cite-this-work">Cite this work</h2>

<div class="language-bibtex highlighter-rouge"><div class="highlight"><pre class="highlight" tabindex="0"><code><table class="rouge-table"><tbody><tr><td class="rouge-gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
</pre></td><td class="rouge-code"><pre><span class="nc">@article</span><span class="p">{</span><span class="nl">ribeiro2024adaptive</span><span class="p">,</span>
  <span class="na">title</span> <span class="p">=</span> <span class="s">{Adaptive Recommendation Under Fairness Constraints}</span><span class="p">,</span>
  <span class="na">author</span> <span class="p">=</span> <span class="s">{Ribeiro, Diogo and Smith, Ada}</span><span class="p">,</span>
  <span class="na">journal</span> <span class="p">=</span> <span class="s">{Journal of Responsible AI}</span><span class="p">,</span>
  <span class="na">year</span> <span class="p">=</span> <span class="s">{2024}</span><span class="p">,</span>
  <span class="na">volume</span> <span class="p">=</span> <span class="s">{12}</span><span class="p">,</span>
  <span class="na">number</span> <span class="p">=</span> <span class="s">{2}</span><span class="p">,</span>
  <span class="na">pages</span> <span class="p">=</span> <span class="s">{45--63}</span><span class="p">,</span>
  <span class="na">doi</span> <span class="p">=</span> <span class="s">{10.1234/jrai.2024.5678}</span>
<span class="p">}</span>
</pre></td></tr></tbody></table></code></pre></div></div>

<p>Add this BibTeX block to your citation manager or to the <code class="language-plaintext highlighter-rouge">CITATION.cff</code> file when you release accompanying code. The DataLog theme handles footnotes, equations, and code blocks seamlessly in a single article.</p>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:1">
      <p>Li, Lihong, et al. "A Contextual-Bandit Approach to Personalized News Article Recommendation." <em>WWW</em> (2010).&nbsp;<a href="#fnref:1" class="reversefootnote" role="doc-backlink">↩</a></p>
    </li>
    <li id="fn:2">
      <p>Zafar, Muhammad Bilal, et al. "Fairness Beyond Disparate Treatment &amp; Disparate Impact." <em>WWW</em> (2017).&nbsp;<a href="#fnref:2" class="reversefootnote" role="doc-backlink">↩</a></p>
    </li>
  </ol>
</div>]]></content><author><name>Diogo Ribeiro</name></author><category term="research" /><category term="publication" /><category term="citations" /><summary type="html"><![CDATA[Publishing reproducible scholarship requires more than compelling charts. This template demonstrates how to structure a research article, cite related work, and provide BibTeX metadata so colleagues can reference your study quickly.]]></summary></entry><entry><title type="html">SQL Optimization Playbook for Warehouse Analysts</title><link href="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/05/sql-optimization-guide/" rel="alternate" type="text/html" title="SQL Optimization Playbook for Warehouse Analysts" /><published>2024-04-05T00:00:00+01:00</published><updated>2024-04-05T00:00:00+01:00</updated><id>https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/05/sql-optimization-guide</id><content type="html" xml:base="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/05/sql-optimization-guide/"><![CDATA[<p>Modern cloud warehouses give you sophisticated tuning knobs, but documentation often lags behind. This guide shows how to annotate SQL plans, highlight critical snippets, and attach performance artifacts so every reviewer can reproduce improvements.</p>

<h2 id="baseline-query">Baseline query</h2>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight" tabindex="0"><code><table class="rouge-table"><tbody><tr><td class="rouge-gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
</pre></td><td class="rouge-code"><pre><span class="k">SELECT</span>
  <span class="n">customer_id</span><span class="p">,</span>
  <span class="k">SUM</span><span class="p">(</span><span class="n">spend</span><span class="p">)</span> <span class="k">AS</span> <span class="n">total_spend</span><span class="p">,</span>
  <span class="k">AVG</span><span class="p">(</span><span class="n">spend</span><span class="p">)</span> <span class="k">AS</span> <span class="n">avg_spend</span><span class="p">,</span>
  <span class="n">ROW_NUMBER</span><span class="p">()</span> <span class="n">OVER</span> <span class="p">(</span><span class="k">PARTITION</span> <span class="k">BY</span> <span class="n">customer_id</span> <span class="k">ORDER</span> <span class="k">BY</span> <span class="n">purchase_ts</span> <span class="k">DESC</span><span class="p">)</span> <span class="k">AS</span> <span class="n">purchase_rank</span>
<span class="k">FROM</span> <span class="n">analytics</span><span class="p">.</span><span class="n">fact_orders</span>
<span class="k">WHERE</span> <span class="n">purchase_ts</span> <span class="o">&gt;=</span> <span class="n">DATEADD</span><span class="p">(</span><span class="s1">'day'</span><span class="p">,</span> <span class="o">-</span><span class="mi">90</span><span class="p">,</span> <span class="k">CURRENT_DATE</span><span class="p">)</span>
<span class="k">GROUP</span> <span class="k">BY</span> <span class="mi">1</span>
</pre></td></tr></tbody></table></code></pre></div></div>

<p>Use the built-in explain tools to gather diagnostics:</p>

<div class="language-sql highlighter-rouge"><div class="highlight"><pre class="highlight" tabindex="0"><code><table class="rouge-table"><tbody><tr><td class="rouge-gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
12
13
</pre></td><td class="rouge-code"><pre><span class="k">EXPLAIN</span> <span class="k">USING</span> <span class="n">JSON</span>
<span class="k">SELECT</span> <span class="o">*</span>
<span class="k">FROM</span> <span class="p">(</span>
  <span class="c1">-- baseline aggregation</span>
  <span class="k">SELECT</span>
    <span class="n">customer_id</span><span class="p">,</span>
    <span class="k">SUM</span><span class="p">(</span><span class="n">spend</span><span class="p">)</span> <span class="k">AS</span> <span class="n">total_spend</span><span class="p">,</span>
    <span class="k">AVG</span><span class="p">(</span><span class="n">spend</span><span class="p">)</span> <span class="k">AS</span> <span class="n">avg_spend</span>
  <span class="k">FROM</span> <span class="n">analytics</span><span class="p">.</span><span class="n">fact_orders</span>
  <span class="k">WHERE</span> <span class="n">purchase_ts</span> <span class="o">&gt;=</span> <span class="n">DATEADD</span><span class="p">(</span><span class="s1">'day'</span><span class="p">,</span> <span class="o">-</span><span class="mi">90</span><span class="p">,</span> <span class="k">CURRENT_DATE</span><span class="p">)</span>
  <span class="k">GROUP</span> <span class="k">BY</span> <span class="mi">1</span>
<span class="p">)</span> <span class="n">src</span>
<span class="k">JOIN</span> <span class="n">analytics</span><span class="p">.</span><span class="n">dim_customer</span> <span class="n">dc</span> <span class="k">USING</span> <span class="p">(</span><span class="n">customer_id</span><span class="p">);</span>
</pre></td></tr></tbody></table></code></pre></div></div>

<h2 id="optimization-checklist">Optimization checklist</h2>

<ol>
  <li>Materialize the 90-day window as an incremental model.</li>
  <li>Cluster the fact table on <code class="language-plaintext highlighter-rouge">purchase_ts</code> to prune partitions.</li>
  <li>Replace repeated JSON parsing with persisted staged columns.</li>
  <li>Cache high-cardinality dimension joins using search optimization services.</li>
</ol>

<blockquote>
  <p><strong>Note:</strong> Store the JSON explain output in <code class="language-plaintext highlighter-rouge">_datasets/</code> so reviewers can diff the query plan over time.</p>
</blockquote>

<h2 id="annotate-performance-wins">Annotate performance wins</h2>

<div class="content-table" role="region" tabindex="0" aria-label="Table 1"><table>
  <thead>
    <tr>
      <th>Change</th>
      <th>Before (s)</th>
      <th>After (s)</th>
      <th>Impact</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Clustering on <code class="language-plaintext highlighter-rouge">purchase_ts</code></td>
      <td>22.4</td>
      <td>9.7</td>
      <td>2.3× faster</td>
    </tr>
    <tr>
      <td>Incremental materialization</td>
      <td>9.7</td>
      <td>4.5</td>
      <td>2.1× faster</td>
    </tr>
    <tr>
      <td>Persisted JSON attributes</td>
      <td>4.5</td>
      <td>3.8</td>
      <td>1.2× faster</td>
    </tr>
  </tbody>
</table></div>

<p>Wrap up by linking to dbt models, scheduling notes, and alert thresholds so stakeholders can keep the warehouse humming.</p>]]></content><author><name>Diogo Ribeiro</name></author><category term="sql" /><category term="performance" /><category term="warehousing" /><summary type="html"><![CDATA[Modern cloud warehouses give you sophisticated tuning knobs, but documentation often lags behind. This guide shows how to annotate SQL plans, highlight critical snippets, and attach performance artifacts so every reviewer can reproduce improvements.]]></summary></entry><entry><title type="html">MLOps Walkthrough with Jupyter Notebook Integration</title><link href="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/04/machine-learning-notebook-integration/" rel="alternate" type="text/html" title="MLOps Walkthrough with Jupyter Notebook Integration" /><published>2024-04-04T00:00:00+01:00</published><updated>2024-04-04T00:00:00+01:00</updated><id>https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/04/machine-learning-notebook-integration</id><content type="html" xml:base="https://diogoribeiro7.github.io/analytics-blog-jekyll/2024/04/04/machine-learning-notebook-integration/"><![CDATA[<p>Production-grade machine learning documentation pairs code, metrics, and narrative. This guide walks through a churn prediction notebook and highlights how the <strong>DataLog</strong> theme embeds notebooks with launch buttons for popular runtimes.</p>

<h2 id="notebook-overview">Notebook overview</h2>

<p>The project notebook <code class="language-plaintext highlighter-rouge">notebooks/churn-segmentation.ipynb</code> contains:</p>

<ul>
  <li>Feature engineering with pandas and scikit-learn <code class="language-plaintext highlighter-rouge">ColumnTransformer</code></li>
  <li>Model training using <code class="language-plaintext highlighter-rouge">xgboost.XGBClassifier</code></li>
  <li>MLflow logging for parameters, metrics, and artifacts</li>
</ul>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight" tabindex="0"><code><table class="rouge-table"><tbody><tr><td class="rouge-gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
</pre></td><td class="rouge-code"><pre><span class="kn">from</span> <span class="n">xgboost</span> <span class="kn">import</span> <span class="n">XGBClassifier</span>
<span class="kn">from</span> <span class="n">sklearn.compose</span> <span class="kn">import</span> <span class="n">ColumnTransformer</span>
<span class="kn">from</span> <span class="n">sklearn.pipeline</span> <span class="kn">import</span> <span class="n">Pipeline</span>
<span class="kn">from</span> <span class="n">sklearn.preprocessing</span> <span class="kn">import</span> <span class="n">OneHotEncoder</span><span class="p">,</span> <span class="n">StandardScaler</span>

<span class="n">numeric</span> <span class="o">=</span> <span class="p">[</span><span class="sh">"</span><span class="s">monthly_charges</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">tenure</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">support_tickets</span><span class="sh">"</span><span class="p">]</span>
<span class="n">categorical</span> <span class="o">=</span> <span class="p">[</span><span class="sh">"</span><span class="s">contract</span><span class="sh">"</span><span class="p">,</span> <span class="sh">"</span><span class="s">region</span><span class="sh">"</span><span class="p">]</span>

<span class="n">preprocess</span> <span class="o">=</span> <span class="nc">ColumnTransformer</span><span class="p">(</span>
    <span class="p">[</span>
        <span class="p">(</span><span class="sh">"</span><span class="s">num</span><span class="sh">"</span><span class="p">,</span> <span class="nc">StandardScaler</span><span class="p">(),</span> <span class="n">numeric</span><span class="p">),</span>
        <span class="p">(</span><span class="sh">"</span><span class="s">cat</span><span class="sh">"</span><span class="p">,</span> <span class="nc">OneHotEncoder</span><span class="p">(</span><span class="n">handle_unknown</span><span class="o">=</span><span class="sh">"</span><span class="s">ignore</span><span class="sh">"</span><span class="p">),</span> <span class="n">categorical</span><span class="p">),</span>
    <span class="p">]</span>
<span class="p">)</span>

<span class="n">model</span> <span class="o">=</span> <span class="nc">Pipeline</span><span class="p">(</span>
    <span class="n">steps</span><span class="o">=</span><span class="p">[</span>
        <span class="p">(</span><span class="sh">"</span><span class="s">preprocess</span><span class="sh">"</span><span class="p">,</span> <span class="n">preprocess</span><span class="p">),</span>
        <span class="p">(</span>
            <span class="sh">"</span><span class="s">classifier</span><span class="sh">"</span><span class="p">,</span>
            <span class="nc">XGBClassifier</span><span class="p">(</span>
                <span class="n">max_depth</span><span class="o">=</span><span class="mi">4</span><span class="p">,</span>
                <span class="n">n_estimators</span><span class="o">=</span><span class="mi">200</span><span class="p">,</span>
                <span class="n">subsample</span><span class="o">=</span><span class="mf">0.8</span><span class="p">,</span>
                <span class="n">colsample_bytree</span><span class="o">=</span><span class="mf">0.9</span><span class="p">,</span>
                <span class="n">eval_metric</span><span class="o">=</span><span class="sh">"</span><span class="s">auc</span><span class="sh">"</span><span class="p">,</span>
            <span class="p">),</span>
        <span class="p">),</span>
    <span class="p">]</span>
<span class="p">)</span>
</pre></td></tr></tbody></table></code></pre></div></div>

<h2 id="launch-options">Launch options</h2>

<div class="cta-group">
  <a class="btn btn--primary" href="https://mybinder.org/v2/gh/DiogoRibeiro7/datalog-notebooks/main?labpath=churn-segmentation.ipynb">Run in Binder</a>
  <a class="btn" href="https://colab.research.google.com/github/DiogoRibeiro7/datalog-notebooks/blob/main/churn-segmentation.ipynb">Open in Colab</a>
  <a class="btn btn--ghost" href="https://github.com/DiogoRibeiro7/datalog-notebooks/blob/main/churn-segmentation.ipynb">View on GitHub</a>
</div>

<p>Readers can open the notebook in the environment of their choice, while the theme preserves accessibility labels for screen readers.</p>

<h2 id="track-experiments">Track experiments</h2>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight" tabindex="0"><code><table class="rouge-table"><tbody><tr><td class="rouge-gutter gl"><pre class="lineno">1
2
3
4
5
6
7
8
</pre></td><td class="rouge-code"><pre><span class="kn">import</span> <span class="n">mlflow</span>

<span class="n">mlflow</span><span class="p">.</span><span class="nf">set_experiment</span><span class="p">(</span><span class="sh">"</span><span class="s">churn-segmentation</span><span class="sh">"</span><span class="p">)</span>
<span class="k">with</span> <span class="n">mlflow</span><span class="p">.</span><span class="nf">start_run</span><span class="p">(</span><span class="n">run_name</span><span class="o">=</span><span class="sh">"</span><span class="s">xgboost-baseline</span><span class="sh">"</span><span class="p">):</span>
    <span class="n">model</span><span class="p">.</span><span class="nf">fit</span><span class="p">(</span><span class="n">train_features</span><span class="p">,</span> <span class="n">train_labels</span><span class="p">)</span>
    <span class="n">auc</span> <span class="o">=</span> <span class="n">model</span><span class="p">.</span><span class="nf">score</span><span class="p">(</span><span class="n">test_features</span><span class="p">,</span> <span class="n">test_labels</span><span class="p">)</span>
    <span class="n">mlflow</span><span class="p">.</span><span class="nf">log_metric</span><span class="p">(</span><span class="sh">"</span><span class="s">test_auc</span><span class="sh">"</span><span class="p">,</span> <span class="n">auc</span><span class="p">)</span>
    <span class="n">mlflow</span><span class="p">.</span><span class="n">xgboost</span><span class="p">.</span><span class="nf">log_model</span><span class="p">(</span><span class="n">model</span><span class="p">.</span><span class="n">named_steps</span><span class="p">[</span><span class="sh">"</span><span class="s">classifier</span><span class="sh">"</span><span class="p">],</span> <span class="sh">"</span><span class="s">model</span><span class="sh">"</span><span class="p">)</span>
</pre></td></tr></tbody></table></code></pre></div></div>

<p>Embed evaluation tables and charts produced by MLflow in the post so stakeholders understand progress:</p>

<div class="content-table" role="region" tabindex="0" aria-label="Table 1"><table>
  <thead>
    <tr>
      <th>Metric</th>
      <th>Value</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Validation AUC</td>
      <td>0.864</td>
    </tr>
    <tr>
      <td>Test AUC</td>
      <td>0.851</td>
    </tr>
    <tr>
      <td>Drift monitor</td>
      <td>Stable</td>
    </tr>
  </tbody>
</table></div>

<h2 id="checklist-before-deployment">Checklist before deployment</h2>

<ul class="task-list">
  <li class="task-list-item"><input type="checkbox" class="task-list-item-checkbox" disabled="disabled" checked="checked">Notebook executed from top to bottom without errors</li>
  <li class="task-list-item"><input type="checkbox" class="task-list-item-checkbox" disabled="disabled" checked="checked">Model registered with reproducible environment metadata</li>
  <li class="task-list-item"><input type="checkbox" class="task-list-item-checkbox" disabled="disabled" checked="checked">Alert thresholds documented for precision/recall trade-offs</li>
</ul>

<p>DataLog’s notebook integration keeps workflows transparent—link to runnable notebooks, surface experiment logs, and capture decisions alongside the code that produced them.</p>]]></content><author><name>Diogo Ribeiro</name></author><category term="machine-learning" /><category term="mlops" /><category term="notebooks" /><summary type="html"><![CDATA[Production-grade machine learning documentation pairs code, metrics, and narrative. This guide walks through a churn prediction notebook and highlights how the DataLog theme embeds notebooks with launch buttons for popular runtimes.]]></summary></entry></feed>