feat: rebuild evidence-first application workflow

This commit is contained in:
2026-07-27 17:56:15 +02:00
parent c24892f381
commit fe5f24704f
57 changed files with 3815 additions and 3555 deletions
@@ -21,14 +21,14 @@ Each achievement has a **Significance** line (why it matters to any reader) and
---
### SW-1: AWS Migration of Legacy ETL Stack
**Significance:** Demonstrates hands-on cloud migration ownership at scale — a tier-1 signal for all data engineering and platform roles. AWS is the market-dominant cloud; owning a full migration from legacy to serverless is a top-of-market achievement.
**Significance:** Demonstrates hands-on migration delivery for pipelines in Dennis's owned domains and contribution to a wider enterprise programme. It is a strong AWS/data-engineering signal without implying ownership of the company-wide migration.
| Role Type | Priority | Lead Verb | Framing Angle |
|-----------|----------|-----------|---------------|
| Staff/Senior Data Engineer | HIGH | Migrated | Lead with scale + operational impact (reduced overhead) |
| Analytics Engineer | HIGH | Migrated | Lead with "enabling analytics outcomes" — tie to downstream stakeholder value |
| ML/AI Engineer | MED | Migrated | Frame as "building the data infrastructure enabling ML workflows" |
| Data Platform/Infra | HIGH | Architected | Lead with cloud-native architecture decisions; de-emphasize analytics framing |
| Data Platform/Infra | HIGH | Migrated / implemented | Lead with verified AWS services and the owned-domain scope; do not claim company-wide architecture ownership |
**Overclaiming warning:** No specific throughput/volume numbers available — do not invent. Use qualitative impact (operational overhead reduction, scalability improvement).
@@ -72,15 +72,15 @@ Each achievement has a **Significance** line (why it matters to any reader) and
---
### SW-5: Security Champion — 3 Consecutive Years
**Significance:** 3 consecutive years = institutional trust, not just a one-time training. Signals security ownership across the DevSecOps lifecycle — rare for a data engineer to hold this level of security designation.
### SW-5: Security Champion — 2025/2026 (team role, NOT an award)
**Significance:** Modest. This is a **rotating team role** (security point of contact), held for **2025/2026 only** — corrected 2026-07-27. It is not an award, not a distinction, and not "3 consecutive years" (an earlier version of this file claimed that; it was wrong).
**DEFAULT: OMIT.** Include only when the JD explicitly requires security or DevSecOps experience. Never place under Awards/Honors.
| Role Type | Priority | Lead Verb | Framing Angle |
|-----------|----------|-----------|---------------|
| Staff/Senior Data Engineer | MED | Designated | Include as breadth signal for senior roles; shows accountability beyond code |
| Analytics Engineer | LOW | — | Omit — not differentiating for this audience |
| ML/AI Engineer | MED | Designated | Include for AI product companies where model security/compliance is relevant |
| Data Platform/Infra | HIGH | Designated | Lead DevSecOps angle — infrastructure roles care about security compliance |
| All role types (JD silent on security) | OMIT | — | Leave it out — it costs a bullet slot and reads as padding |
| Any role type (JD explicitly requires security/DevSecOps) | MED | Serve as | State plainly as a team role for 2025/2026; cite the 100h training + assessment, claim no more |
---
@@ -298,7 +298,7 @@ Each achievement has a **Significance** line (why it matters to any reader) and
| SW-2 Component Owner | HIGH | HIGH | MED | HIGH |
| SW-3 K8s + GitLab | HIGH | MED | HIGH | HIGH |
| SW-4 B2B Products | MED | HIGH | LOW | LOW |
| SW-5 Security Champion | MED | LOW | MED | HIGH |
| SW-5 Security Champion | LOW | LOW | LOW | LOW |
| SW-6 PySpark | MED | LOW | MED | MED |
| BS-1 ML Inference | HIGH | LOW | HIGH | HIGH |
| BS-2 Data Services | HIGH | MED | MED | HIGH |
+35 -124
View File
@@ -1,133 +1,44 @@
# AI Fingerprint Avoidance Rules
# Authenticity and Natural-Writing Rules
> **Architecture note:** The primary defense against AI detection is the generation protocol — specific facts from experience files, char limits, JD-specific vocabulary, named entities. This file is a secondary safety net for word/phrase/structural patterns.
> The primary defense against generic or suspicious application writing is verifiable specificity.
> There is no reliable punctuation checklist for determining whether text was written with AI.
---
## Mandatory Authenticity Checks
## 1. Banned Words
1. Trace every experience claim to `resume_builder/canonical/claims.json` and an experience file.
2. Preserve ownership scope and allowed verbs. Never intensify a verb merely to match the JD.
3. Use the JD's terminology only when it accurately names the demonstrated work.
4. Do not copy long phrases from the JD or mirror its requirement order mechanically.
5. Do not add a tool, scale, customer, metric or outcome because it is plausible.
6. Keep bridges explicit: say the demonstrated technology first, then explain transferability if useful.
7. Prefer a concrete fact over a generic adjective.
8. A claim must be answerable in an interview with a specific example.
**Tier 1 — Dead Giveaways (NEVER use in any output):**
delve, tapestry, multifaceted, pivotal, realm, synergy, paradigm, holistic, nuanced, foster, embark, leverage (as verb), utilize, harness, spearhead, cornerstone, landscape (metaphorical), journey (metaphorical), cutting-edge, novel, innovative (unless quoting a JD), groundbreaking
## Natural Resume Writing
**Banned Adjectives (use replacement):**
- Start bullets with the work or result, not a framing phrase.
- Mix short and longer bullets according to information content; identical lengths are not a goal.
- Use ordinary action verbs. Repetition is acceptable when the same verb is accurate.
- Avoid empty claims such as innovative, cutting-edge, proven track record, uniquely positioned or passionate about.
- Use metrics only when verified. Qualitative outcomes are acceptable when their source is clear.
- Keep the employer, formal title, dates and location more prominent than tailored narrative language.
- Do not use first person in the resume. First person is appropriate in a cover letter.
| Banned | Replacement |
|--------|-------------|
| robust | strong, reliable |
| comprehensive | thorough, broad |
| innovative | new, original (or omit) |
| pivotal | key, central |
| meticulous | careful, precise |
| diverse | varied, wide-ranging |
| extensive | broad, deep, 10+ years of |
## Natural Cover-Letter Writing
**Banned Verbs (use replacement):**
- Write a letter only when it adds value or the employer requests it.
- Open with the candidate-role connection; company research should support that connection, not dominate it.
- Use one or two specific reasons for interest, not a paragraph of company-news paraphrase.
- Do not repeat resume bullets. Explain a decision, motivation, transition or working style the resume cannot show.
- Avoid defensive gap lists. A necessary bridge should be short and evidence-led.
- Contractions and normal punctuation are allowed. There is no em-dash quota and no ban on sentences ending in an -ing word.
| Banned | Replacement |
|--------|-------------|
| leverage | use, apply, draw on |
| utilize | use |
| harness | apply, use, draw on |
| spearhead | lead, start, launch |
| foster | support, build, grow |
| facilitate | run, lead, coordinate, enable |
| showcase | show, demonstrate |
| underscore | show, highlight |
| bolster | strengthen, support |
## Post-Generation Checklist
**Banned Adverbs:** meticulously, notably, subsequently (use "then" or "later"), remarkably, seamlessly, thereby
**Banned Nouns (metaphorical use):** tapestry, landscape, journey, realm, synergy, paradigm, cornerstone
**Technical exceptions:** "landscape" is fine when literal (e.g., "free energy landscape," "threat landscape"). "Novel" is fine when quoting a JD verbatim. Judge by context.
---
## 2. Banned Phrases
**Opening / transition phrases:**
- "In today's rapidly evolving..."
- "At the forefront of..."
- "It is worth noting that..."
- "This experience has taught me..."
- "I am uniquely positioned to..."
- "In an era of..."
**Resume / CL specific:**
- "proven track record"
- "passionate about" (use specific interest instead)
- "I am excited to apply" (use concrete reason instead)
- "demonstrated ability to" (just state what you did)
- "strong foundation in"
- "well-versed in"
- "adept at"
**Academic / research:**
- "groundbreaking research"
- "cutting-edge methodology"
- "novel approach" (say what is new about it)
- "significant contributions to the field"
- "at the intersection of X and Y" (name the specific intersection)
---
## 3. Structural Rules
### Sentence-Level
- **No reframe pattern:** Never use "It's not X — it's Y" constructions
- **No rhetorical Q+A:** Never ask a question then answer it ("What makes this unique? The answer is...")
- **No gerund fragment stacking:** Avoid sequences of 3+ "-ing" phrases ("developing, testing, and deploying...")
- **No -ing analysis endings on bullets:** This is the **#1 structural AI marker**. Bullets must NOT end with "-ing" phrases like "...advancing the field," "...contributing to improved Y," "...enabling new Z." Fix: restructure so the bullet ends with a concrete result, metric, or object. Example: "...contributing to a 15% reduction" is fine (ends with metric); "...contributing to improved efficiency" is not (vague -ing ending).
- **Max 2 em-dashes per document:** Count all `---` in the full .tex file (resume or CL). If more than 2, replace extras with commas, semicolons, or parentheses. Fellowships/Honors items use `. ` not `---`.
- **Post-gen scan:** After generating any document, scan all bullets for -ing endings. Flag and fix any found.
### Prose-Level
- **Vary sentence length:** Mix short (8-12 words) with long (20-30 words). Three consecutive same-length sentences flag as AI.
- **No same-structure paragraph starts:** If P1 opens "My research...", P2 must NOT open "My experience..." P3 must NOT open "My approach..."
- **No constant triplet structures:** Avoid "X, Y, and Z" in more than 2 sentences per document. Use pairs, single items, or lists of 4+.
---
## 4. Positive Markers (signals of human writing)
1. **Specific details:** "Ran 847 MD simulations on protein variants" not "Conducted extensive simulations"
2. **Front-loaded specifics:** Lead with the concrete thing, not the framing
3. **Named entities:** Tool names, method names, journal names, institution names
4. **Audience-appropriate jargon:** Use the JD's vocabulary, not generic synonyms
5. **Short connecting words:** "so," "but," "and," "then" — not "consequently," "however," "additionally," "subsequently"
6. **First-person specificity in CLs:** "I built" not "Was responsible for building"
7. **Inside knowledge:** Reference specific group names, facility names, programmatic areas
8. **Sentence length variety:** Deliberate mix of 8-word and 25-word sentences
9. **Occasional "And"/"But" sentence openers** in CLs (1-2 per page max)
10. **Contractions in CLs:** "I've" and "didn't" are acceptable in industry CLs (not academic)
11. **One human detail per CL page:** A specific lab memory, a conference conversation, a problem that kept you up — concrete and brief
---
## 5. CL-Specific Note
Cover letters are the most vulnerable document to AI detection because they are prose-heavy and readers have strong intuitions about "how people write." All rules above apply with extra weight in CLs. Pay special attention to:
- Opening sentence (must be specific to the company, not generic)
- Sentence length variety (CLs with uniform 15-20 word sentences read as AI)
- Em-dash usage (CLs accumulate em-dashes fastest — max 2 for the entire letter)
---
## 6. Post-Generation Critique Scan Checklist
Run this 12-item scan on every generated document before presenting to the user:
1. [ ] Any Tier 1 banned word present? (Search for each)
2. [ ] Any banned phrase from Section 2?
3. [ ] More than 2 em-dashes (`---`) in the document?
4. [ ] Any bullet ending with an -ing analysis phrase?
5. [ ] Three or more consecutive sentences of similar length?
6. [ ] Paragraph starts repeat the same structure (e.g., "My research...", "My experience...")?
7. [ ] More than 2 "X, Y, and Z" triplet structures in the document?
8. [ ] CL opens with a generic phrase instead of a company-specific reference?
9. [ ] Any metaphorical use of "landscape," "journey," "realm," or "tapestry"?
10. [ ] Passive voice in more than 20% of bullet verbs?
11. [ ] Fellowships/Honors items use `---` instead of `. `?
12. [ ] Any adverb from the banned list (meticulously, notably, subsequently, etc.)?
**If any item fails:** Fix before presenting. These are not optional polish — they are detectable AI patterns.
- [ ] Every claim passes the canonical validator.
- [ ] Every listed skill has an interview-ready evidence example.
- [ ] No sentence converts company or industry scale into personal impact.
- [ ] No marketing headline is presented as a historical job title or customer-delivery fact.
- [ ] No unsupported metric or causal result appears.
- [ ] The resume sounds like one technically precise person, not a rearranged JD.
- [ ] The cover letter is optional by policy and adds information beyond the resume.
+10 -10
View File
@@ -1,7 +1,7 @@
# Significance Research: Bosch Semiconductor — Data Analysis Engineer
> Use in cover letters and summaries — NOT in resume bullet text.
> Particularly valuable for semiconductor industry JDs.
> Optional semiconductor context only. Reverify external claims before use.
> Never convert generic fab scale, yield economics or industry trends into Dennis's personal impact.
---
@@ -16,19 +16,19 @@
- Offline ML analysis (batch — not real-time, misses process drifts)
- Inline ML inference (real-time, containerized — current best practice)
**Why Dennis's experience matters:** Deploying ML inference into a 24/7 fab is operationally much harder than deploying to a web server. There are no maintenance windows, hardware is constrained, and a model failure affects production throughput. Dennis designed and executed the integration strategy for this environment — a level of MLOps maturity that few data engineers have encountered.
**Why Dennis's experience matters:** Dennis designed and executed an integration strategy for containerized ML inference in a continuously operating fab environment. Do not add claims about maintenance windows, hardware constraints, throughput impact or rarity unless they are verified for his system.
**Differentiation:** The combination of Docker containerization + Kubernetes orchestration + Ansible automation in a 24/7 constrained environment is a rare and credible production ML deployment signal.
**Differentiation:** Docker, Kubernetes and Ansible used for production inference integration provide direct deployment evidence. This supports ML-platform/MLOps positioning without implying model-development ownership.
---
### Semiconductor Data Domains — Field Context
**Defect Management:**
Semiconductor defect management involves tracking, classifying, and correlating defects found during inline inspection (optical, SEM) and end-of-line electrical test. Key data challenges: high-dimensional spatial data (wafer maps), multi-step process correlation, and connecting defect signatures to root causes (process excursions, equipment issues). Dennis built data pipelines and ML systems directly in this domain.
Semiconductor defect management provides the domain context for Dennis's work. His verified contributions cover data services, analytics applications, wafer-map visualizations and ML-inference integration; do not generalize this into ownership of every defect-management pipeline or ML system.
**Semiconductor Parameter Testing:**
Parametric testing measures electrical characteristics (threshold voltages, leakage currents, resistance) of test structures on each wafer. The data volume is massive — hundreds of parameters across thousands of dies per wafer, across thousands of wafers per day. Data engineering for parametric test requires efficient storage, fast query access, and statistical analysis capabilities. Dennis built data services that fed parametric testing analysis teams.
Dennis built data services for semiconductor analysis teams in the parameter-testing domain. Generic wafer or data-volume figures must not be presented as the scale of his systems without direct evidence.
**Process Analysis:**
Process analysis correlates equipment parameters (temperature, pressure, gas flow) with downstream wafer yield and defect outcomes. This is the domain where data engineering meets process engineering — the pipelines must be reliable and the data must be accurate, because process decisions (equipment maintenance, recipe adjustments) depend on it.
@@ -40,13 +40,13 @@ Process analysis correlates equipment parameters (temperature, pressure, gas flo
### Field Overview: Data & AI in Semiconductor Manufacturing (20242026)
The semiconductor industry is undergoing a major digital transformation driven by:
1. **Process complexity:** 300mm fabs with 1000+ process steps generate petabytes of data; manual analysis can no longer keep pace
2. **Yield pressure:** At leading-edge nodes, even 1% yield improvement has enormous economic value — data-driven yield optimization is a strategic priority
1. **Process complexity:** 300mm semiconductor production creates complex data and operational requirements; do not attach generic petabyte or process-step figures to Dennis's work
2. **Yield and quality:** Data-driven process and defect analysis matter commercially, but Dennis has no verified personal yield-improvement metric
3. **AI/ML adoption:** Computer vision for inline inspection, predictive maintenance for equipment, and ML-based process optimization are all actively deployed at tier-1 fabs (TSMC, Intel, Samsung)
4. **Talent scarcity:** Candidates who combine data engineering depth with semiconductor domain knowledge are extremely rare — most data engineers lack the domain; most process engineers lack the data skills
4. **Candidate distinction:** Dennis combines production data engineering with direct semiconductor-fab application experience; avoid unsourced scarcity claims
**Target companies for semiconductor JDs:**
ASML, Infineon, GlobalFoundries, ams OSRAM, Microchip Technology, ON Semiconductor, Renesas, NXP, STMicroelectronics, Bosch (again), TSMC (Europe fabs in Dresden area), Wolfspeed, SiCrystal, Elmos
**CL hook for semiconductor JDs:**
> "Semiconductor manufacturing analytics is one of the most data-intensive and operationally demanding domains in industry. At Bosch Semiconductor in Dresden, I worked directly in the data domains that matter most — Defect Management, Semiconductor Parameter Testing, and Process Analysis — building the pipelines and analytics platforms that engineers relied on for real-time production decisions. That domain knowledge, combined with my experience deploying ML-based defect classification into a 24/7 fab, is what I'd bring to [Company]."
> "At Bosch Semiconductor in Dresden, I developed data services and analytics applications for defect-management and process-analysis teams, and integrated containerized ML inference into a 24/7 fab environment. That combination of domain familiarity and production delivery is what I would bring to [Company]."
+12 -12
View File
@@ -1,7 +1,7 @@
# Significance Research: Swisscom — Staff Data, Analytics & AI Engineer
> Use in cover letters and summaries — NOT in resume bullet text.
> Provides field context demonstrating awareness of the data engineering landscape.
> Optional context only. Reverify every market/company statement from a current primary source before use.
> Never convert company scale, industry trends or likely benefits into Dennis's personal impact.
---
@@ -9,24 +9,24 @@
**The problem:** Legacy enterprise data warehouses (Teradata, Oracle) are expensive to scale, inflexible for modern analytics workloads, and difficult to integrate with ML pipelines. The industry-wide shift to cloud-native data platforms (AWS, Azure, GCP) is driven by cost, elasticity, and the rise of the data lakehouse pattern.
**Competing approaches:** Most enterprises face a choice between lift-and-shift (rehosting on cloud VMs — minimal benefit), re-platforming (moving to managed services like Redshift), or full re-architecture to a lakehouse (S3 + Athena/Iceberg + Glue). The lakehouse pattern (Apache Iceberg on S3 + Athena) is increasingly the de facto standard for cost-efficient, ACID-compliant analytics at scale.
**Architecture context:** Enterprise migrations may involve rehosting, re-platforming or re-architecture. Use this only to explain the technical choices in a verified job-specific context; do not assert a universal standard.
**Why this matters:** Swisscom serves millions of Swiss customers across mobile, broadband, and enterprise — the data volume is significant. Moving Fulfillment data pipelines to a cloud-native architecture directly affects the speed and cost of analytics for business-critical processes.
**Why this matters:** Fulfillment pipelines are business-critical. No personal data-volume, cost-saving or speed-improvement metric is currently verified.
**Differentiation:** Dennis didn't just configure existing pipelines in a new environment — he introduced Apache Iceberg (open table format with time-travel and schema evolution), AWS Glue Tables as the catalog, and CloudFormation for IaC provisioning. This reflects current best practices in data lakehouse architecture, not a basic ETL migration.
**Candidate-specific evidence:** Dennis used Iceberg, AWS Glue/Athena and CloudFormation while migrating pipelines in his owned domains. Do not claim he selected these technologies for Swisscom or introduced them company-wide unless separately verified.
**Field overview: Data Lakehouse Architecture (20242026)**
The data lakehouse pattern — combining the scalability of data lakes (S3, ADLS) with the ACID guarantees and query performance of data warehouses — has become the dominant architecture for new data platform builds. Apache Iceberg has emerged as the leading open table format, supported by AWS (Athena, Glue), Databricks (as Delta Lake alternative), and Snowflake. Engineers who have implemented Iceberg in production (not just read about it) are in high demand as organizations migrate off proprietary DWH systems.
Open table formats such as Apache Iceberg are relevant context for lakehouse roles. Reverify any market-share or "dominant architecture" statement from current primary sources before using it in a cover letter. Context must never be converted into a claim about Dennis's personal system scale or architecture authority.
---
### SW-2: Component Ownership at Scale — Field Context
**The problem:** In large data engineering teams at enterprise companies, the "Component Owner" model is how organizations assign accountability for production systems. Unlike a ticket-based dev model, Component Owners are responsible for a system's full lifecycle: reliability, compliance, SLA, on-call, and stakeholder communication. This is a Staff-engineer-level responsibility.
**The evidence:** At Swisscom, Dennis's Component Owner role covers production operation, data quality, governance, incidents and on-call obligations for Fulfillment ETL pipelines. Describe those verified responsibilities directly; do not use a generic title-equivalence claim.
**Why this matters:** Swisscom's Fulfillment domain carries business-critical data — provisioning, activating, and tracking customer service orders for Switzerland's largest telecom. Pipeline failures in this domain directly impact customer experience and revenue.
**Why this matters:** Swisscom's Fulfillment domain carries business-critical operational data. Avoid claiming a quantified customer or revenue effect without evidence.
**Differentiation:** Dennis holds this responsibility as a Staff Engineer (Engineer IV) — the same person building the pipelines is accountable for their reliability in production. This is the "full-stack data engineer" model that platform teams increasingly demand.
**Candidate-specific value:** Dennis combines implementation with production accountability for his components. That is the defensible distinction; broader market-demand claims require fresh sourcing.
---
@@ -34,7 +34,7 @@ The data lakehouse pattern — combining the scalability of data lakes (S3, ADLS
**The problem:** Data pipelines have traditionally been deployed on bare metal or VMs, leading to environment inconsistency, difficult scaling, and slow deployments. The shift to Kubernetes for data workloads (not just web services) reflects the maturation of the data platform discipline.
**Industry trend:** Running data applications (Airflow, Spark, custom Python pipelines) on Kubernetes is now standard practice at mature data organizations. GitLab CI/CD with Kubernetes deployment is the Swiss/European enterprise standard (as opposed to GitHub Actions + AWS ECS in US-heavy startups).
**Industry context:** Kubernetes and CI/CD are recognizable production-delivery signals. Do not claim a regional or industry standard without current sourcing.
**Differentiation:** Swisscom's use of Kubernetes for Python data applications confirms production-grade container orchestration for data workloads — not just a dev/test environment.
@@ -44,8 +44,8 @@ The data lakehouse pattern — combining the scalability of data lakes (S3, ADLS
The data engineering discipline has undergone a significant shift in the past 3 years:
1. **From batch to streaming:** Kafka-based event-driven architectures have replaced many nightly batch processes
2. **From proprietary DWH to open lakehouse:** Teradata/Oracle S3 + Athena/Iceberg is the dominant migration pattern
2. **From proprietary DWH to open lakehouse:** Dennis has direct experience moving owned-domain pipelines from Teradata/Oracle processing to S3 + Athena/Iceberg within a wider programme
3. **From manual to automated infra:** CloudFormation, Terraform, and Pulumi have made IaC standard for data platform teams
4. **From separated to embedded ML:** Data engineers who can own the ML data layer (not just supply data to a separate ML team) are increasingly valuable
Dennis's current stack (Kafka, PySpark, AWS S3/Glue/Athena/Iceberg, Kubernetes, GitLab CI/CD, CloudFormation) maps precisely to this modern paradigm.
Dennis's current stack includes Kafka, PySpark, AWS S3/Glue/Athena/Iceberg, Kubernetes, GitLab CI/CD and CloudFormation. Use the named evidence; avoid generic claims that it maps "precisely" to every target platform.
+86 -174
View File
@@ -1,195 +1,107 @@
# Skills Taxonomy — Dennis Thiessen
# Skills Taxonomy — Evidence-First
> Generated: 2026-03-28
> Sources: All 10 extractions + 6 experience files
> Use this file when populating the Technical Skills section of resume/CV.
> Canonical authority: `resume_builder/canonical/claims.json`.
> This file helps select and group skills; it may not promote a skill beyond its canonical evidence level.
---
## Evidence Levels
## Summary Stats
| Level | Meaning | Output rule |
|---|---|---|
| Production — current | Used in current professional delivery | May appear plainly when relevant |
| Production — historical | Shipped professionally, but not current | Include with role/date context when recency matters |
| Hands-on — current | Used directly, but without verified production ownership | Use precise verbs such as used, configured or integrated |
| Project / proof of concept | Used in a bounded PoC | Label the PoC; never imply platform-scale operation |
| Certification / coursework | Learned through formal study | Keep in certification context; not professional experience |
| Unverified / never used | No reliable evidence | Do not include |
- **Total unique skills:** 65+
- **Proficiency levels:** Expert (daily use, owned systems) | Proficient (shipped work, comfortable teaching) | Familiar (used in project, not current)
- **Certification-backed skills:** AWS (SAA cert + Udacity DataEng), Software Architecture (iSAQB), AI/ML (Udacity AI for Trading, IBM AI Engineering)
Do not use Expert/Proficient/Familiar labels in resumes. Evidence and recency are more useful than self-ratings.
---
## Current Production Core
## Category 1: Programming Languages
| Skill | Evidence | Typical use |
|---|---|---|
| Python | Swisscom pipelines/apps; prior Bosch/Vizrt work | Always for data/platform roles |
| SQL | Swisscom and prior data roles | Always for data roles |
| PySpark | Swisscom current work | When distributed processing is relevant |
| Apache Kafka | Swisscom production ingestion | Data/event-driven roles |
| Apache Airflow | Swisscom AWS migration scope | Orchestration/data roles |
| AWS | Swisscom production work; SAA certification | AWS-relevant roles |
| S3, Glue, Athena, Iceberg, Redshift | Swisscom owned-domain migration/data products | Name only relevant services |
| CloudFormation / IaC | Swisscom production provisioning | Say CloudFormation; never substitute Terraform |
| Kubernetes, Docker | Swisscom application delivery; Bosch ML integration | Production platform/MLOps roles |
| GitLab CI/CD | Swisscom delivery | Platform and engineering roles |
| Oracle, Teradata | Swisscom pipelines | Data roles when relevant |
| Skill | Proficiency | Evidence | Resume Weight |
|-------|-------------|----------|---------------|
| Python | Expert | Swisscom (pipelines, apps), Bosch (data services), Fraunhofer (ML/NLP), Vizrt (backend + tests) | HIGH |
| SQL (multi-dialect) | Expert | All positions — Oracle, Impala, Teradata, MS SQL, Postgres, MySQL | HIGH |
| PySpark | Proficient | Swisscom Staff level (LinkedIn confirmed) | HIGH |
| Java | Proficient | Fraunhofer (SCEDAS, MISSION), Bosch (data services), Generali (J2EE), Capgemini | MED |
| C# | Proficient | Bosch (data services, Spotfire extensions), Fraunhofer (SCEDAS) | MED |
| JavaScript / TypeScript | Proficient | Fraunhofer (MISSION, Express.js), CV skills list | MED |
| C++ | Proficient | Vizrt (backend transcoding), Generali (CV) | LOW |
| VBA | Familiar | Student assistant role (Bundeswehr Uni, 2013) — very minor | LOW |
## Historical Production Evidence
---
| Skill | Evidence | Constraint |
|---|---|---|
| Java | Bosch, Fraunhofer, Generali | Historical; do not imply current daily use |
| C# | Bosch and Fraunhofer | Historical; strong when Spotfire/.NET is relevant |
| C++ | Vizrt distributed backend | Limited historical evidence |
| JavaScript / Express.js | Fraunhofer MISSION | Historical and bounded; TypeScript is unverified |
| Hadoop / Impala | Bosch data services | Historical production context |
| Ansible | Bosch ML integration | Historical production context |
| Jenkins | Fraunhofer and Generali | Historical production context |
| BDD, Selenium, JBehave | Generali | Earlier-career testing evidence |
| TIBCO Spotfire | Bosch co-ownership and C# extensions | Preserve co-ownership |
## Category 2: Data Engineering & Pipelines
## ML, AI and Observability
| Skill | Proficiency | Evidence | Resume Weight |
|-------|-------------|----------|---------------|
| ETL/ELT design & operation | Expert | Swisscom (component owner), Bosch (data services) | HIGH |
| Apache Kafka | Expert | Swisscom (ingestion pipelines), Bosch (ELK PoC) | HIGH |
| Apache Airflow | Proficient | Swisscom (AWS migration stack) | HIGH |
| SAP BODS | Proficient | Swisscom (legacy ETL) | MED |
| Teradata DWH | Proficient | Swisscom (DWH architecture + operation) | MED |
| Hadoop / ImpalaSQL | Proficient | Bosch (data services over Hadoop) | MED |
| Data modeling | Proficient | Swisscom (data products), Bosch (pipeline design) | MED |
| SQL performance tuning | Proficient | CV (explain plans, indexes, partitions) | MED |
| Apache Spark / PySpark | Proficient | Swisscom (big data processing) | HIGH |
| dbt | Not confirmed | Not in any extraction — do not claim | — |
| Skill | Evidence level | Safe framing |
|---|---|---|
| ML inference deployment | Production — historical | Integrated containerized inference into a 24/7 Bosch fab |
| Image classification | Production application context | Worked on inference integration; model-training ownership not verified |
| MLOps | Bounded production evidence | Use only when defined as deployment/operation, not full model lifecycle |
| NLP / speech recognition | Research-project contribution | Contributed components at Fraunhofer; no publication/model ownership |
| ELK, Kafka anomaly detection | Proof of concept | Always retain the PoC label |
| Grafana, Prometheus, Loki | Proof-of-concept/monitoring context | Do not imply enterprise observability ownership |
| LiteLLM | Hands-on — current | LLM API gateway use/integration; no serving-platform ownership |
| Domain-grounded assistants/custom GPTs | Hands-on — current | Configured with curated knowledge; no fine-tuning or formal evaluation |
| Copilot, Kiro | Hands-on — current | AI-assisted engineering tools, not LLM product engineering |
---
## Certification-Only Signals
## Category 3: Cloud & Infrastructure
| Skill | Evidence |
|---|---|
| TensorFlow / Keras | IBM AI Engineering coursework |
| PyTorch | Coursework/personal evidence only; verify before listing outside certification context |
| Spark ML | Coursework context only unless professional evidence is added |
| AI for Trading / quantitative ML | Udacity/WorldQuant Nanodegree |
| Skill | Proficiency | Evidence | Resume Weight |
|-------|-------------|----------|---------------|
| AWS (overall) | Proficient | Swisscom (migration), AWS SAA cert (2024), Udacity DataEng cert (2026) | HIGH |
| AWS S3 | Proficient | Swisscom AWS migration | HIGH |
| AWS Glue | Proficient | Swisscom AWS migration | HIGH |
| AWS Athena | Proficient | Swisscom AWS migration (with Apache Iceberg table format) | HIGH |
| AWS Glue (Jobs + Tables) | Proficient | Swisscom — Glue jobs for ETL + Glue Data Catalog / Glue Tables | HIGH |
| Apache Iceberg | Proficient | Swisscom — S3 + Athena with Iceberg table format (open table format, time-travel, schema evolution) | HIGH |
| AWS Redshift | Proficient | Swisscom AWS migration | HIGH |
| AWS Lambda | Proficient | Swisscom AWS migration | MED |
| AWS Step Functions | Proficient | Swisscom AWS migration | MED |
| AWS CloudFormation | Proficient | Swisscom — IaaS, infrastructure provisioning as code | HIGH |
| Kubernetes (K8s) | Expert | Swisscom (Python app deployment), Bosch (ML inference orchestration) | HIGH |
| Docker | Expert | Bosch (ML containerization, ELK PoC), Fraunhofer (MISSION), Swisscom | HIGH |
| Ansible | Proficient | Bosch (ML orchestration) | MED |
| GitLab CI/CD | Proficient | Swisscom (confirmed Zeugnis) | HIGH |
| Jenkins | Proficient | Fraunhofer (independently set up), Generali (BDD build jobs) | MED |
| CI/CD (general) | Expert | Swisscom, Fraunhofer, Vizrt, Generali — cross-position | HIGH |
| IaC (Infrastructure as Code) | Proficient | Swisscom — AWS CloudFormation confirmed by user | HIGH |
| DevSecOps | Proficient | Swisscom Security Champion ×3 (20232026), 100h training | MED |
## Forbidden Until New Evidence Is Added
---
- LangChain, LangGraph, LlamaIndex, AutoGen, CrewAI or Semantic Kernel
- Azure, Azure ML, Azure OpenAI or AKS
- GCP, BigQuery, Dataflow or Flume hands-on experience
- Terraform
- Formal model or LLM evaluation
- LLM fine-tuning, red-teaming or model-training ownership
- FastAPI, Flask or Django
- TypeScript
- Petabyte-scale ownership
## Category 4: Databases & Storage
## Certifications
| Skill | Proficiency | Evidence | Resume Weight |
|-------|-------------|----------|---------------|
| Oracle DB | Expert | Swisscom (Fulfillment pipelines), Bosch (data services), Generali (web portal) | HIGH |
| Teradata | Proficient | Swisscom (DWH target, architecture) | MED |
| MS SQL Server | Proficient | Fraunhofer (SCEDAS — Entity Framework) | LOW |
| PostgreSQL | Familiar | CV skills list | LOW |
| MySQL | Familiar | CV skills list, RiskAhead project | LOW |
| SQLite | Familiar | Fraunhofer (MISSION microservices) | LOW |
| Hadoop / Impala | Proficient | Bosch (ImpalaSQL data services) | MED |
| Certification | Issuer | Year/status |
|---|---|---|
| AWS Certified Solutions Architect — Associate | AWS | 2024; active to Sep 2027 |
| Data Engineering with AWS Nanodegree | Udacity | 2026 |
| iSAQB CPSA — Foundation | iSAQB | 2016; no expiry |
| ITIL Foundation | PEOPLECERT / AXELOS | 2016; no expiry |
| AI for Trading Nanodegree | Udacity / WorldQuant | 2021 |
| IBM AI Engineering Specialization | IBM / Coursera | Completion year not recorded |
---
The Swisscom Security Champion assignment is not a certification and does not belong in this table.
## Category 5: ML & AI
## Resume Grouping
| Skill | Proficiency | Evidence | Resume Weight |
|-------|-------------|----------|---------------|
| ML inference deployment | Proficient | Bosch (Docker/K8s in 24/7 fab — primary responsibility) | HIGH |
| Image classification | Proficient | Bosch (automated quality monitoring in semiconductor fab) | MED |
| NLP / Speech recognition | Familiar | Fraunhofer ARTUS research project (contributing role) | MED |
| PyTorch | Familiar | CV skills list | LOW |
| Scikit-learn | Familiar | CV skills list | LOW |
| Pandas / NumPy | Proficient | CV (data analysis, pipeline work) | MED |
| Matplotlib / Plotly | Proficient | CV (data visualization, dashboards) | LOW |
| MLOps (general) | Proficient | Bosch (full ML lifecycle: containerize → deploy → monitor in production) | HIGH |
| AI for Trading / Quant ML | Familiar | Udacity AI for Trading Nanodegree (2021) — personal study, not professional | LOW |
| TensorFlow / Keras | Familiar | IBM AI Engineering Specialization (Coursera) | LOW |
| Apache Spark ML | Familiar | IBM AI Engineering (Spark ML course) | LOW |
Use 4--6 compact lines, selected for the JD. A normal International Tech grouping is:
**Proficiency note:** For ML/AI roles, frame Bosch ML deployment as primary evidence. NLP/ARTUS and the Udacity/IBM certs as supporting signals. Do not overstate ML modeling depth — the core strength is ML *infrastructure and deployment*, not research.
1. Languages: Python, SQL; selected historical languages only when required.
2. Data: Kafka, Airflow, PySpark, Oracle/Teradata, data products and governance.
3. Cloud/platform: AWS services, CloudFormation, Kubernetes, Docker, GitLab CI/CD.
4. ML/operations: ML inference deployment and bounded observability evidence.
5. Certifications: one line, only the most relevant credentials.
---
## Category 6: Testing & Quality Engineering
| Skill | Proficiency | Evidence | Resume Weight |
|-------|-------------|----------|---------------|
| Test automation | Expert | Capgemini, Generali, Vizrt — consistent across 3 positions | MED (earlier career) |
| BDD (Behaviour-Driven Development) | Proficient | Generali — introduced PoC, held technical ownership | MED |
| Serenity-BDD / JBehave | Proficient | Generali (confirmed Zeugnis) | LOW |
| Selenium | Proficient | Generali (UI test automation) | LOW |
| pytest | Proficient | CV skills list | MED |
| TDD | Proficient | Capgemini, Generali (confirmed) | LOW |
| HP Quality Center / ALM | Familiar | Capgemini (Zeugnis confirmed) | LOW |
| UIPath RPA | Familiar | Generali (POC developer, confirmed Zeugnis + LinkedIn) | LOW |
| Camunda BPMN | Familiar | Generali (LinkedIn confirmed) | LOW |
| Quality gates (CI/CD) | Proficient | Vizrt (CI/CD integration), Fraunhofer (Jenkins quality gates) | MED |
---
## Category 7: Observability, Monitoring & DevOps Tooling
| Skill | Proficiency | Evidence | Resume Weight |
|-------|-------------|----------|---------------|
| ELK Stack (Elasticsearch/Logstash/Kibana) | Proficient | Bosch (anomaly detection PoC — primary developer) | MED |
| Grafana | Proficient | Bosch (monitoring dashboards) | MED |
| Prometheus | Proficient | Bosch (metrics) | MED |
| Loki | Familiar | Bosch (log aggregation, part of PoC) | LOW |
| Git | Expert | All positions | HIGH |
| Agile / Scrum | Proficient | Swisscom (confirmed Zeugnis — backlog, sprint planning, Product Owner collaboration) | MED |
| Tibco Spotfire | Familiar | Bosch (C# extensions, LinkedIn confirmed) | LOW |
---
## Category 8: Frameworks & APIs
| Skill | Proficiency | Evidence | Resume Weight |
|-------|-------------|----------|---------------|
| Flask / FastAPI / Django | Proficient | CV skills list | MED |
| Express.js | Familiar | Fraunhofer MISSION (microservices) | LOW |
| Entity Framework (.NET) | Proficient | Fraunhofer SCEDAS | LOW |
| Spring Boot | Familiar | Generali (Dispatcher PoC, Apache Camel) | LOW |
| Apache Camel | Familiar | Generali (Dispatcher PoC) | LOW |
| SQLAlchemy | Familiar | CV skills list | LOW |
| Swagger / OpenAPI | Familiar | CV skills list | LOW |
---
## Category 9: Domain Knowledge
| Domain | Depth | Source | Resume Weight |
|--------|-------|--------|---------------|
| Telecom / Enterprise data platforms | Proficient | Swisscom (2+ years, current) | HIGH |
| Semiconductor manufacturing / Industry 4.0 | Proficient | Bosch (3 years) — data domains: Defect Management, Semiconductor Parameter Testing, Process Analysis, Image-based Quality Inspection | MED |
| Maritime logistics | Familiar | Fraunhofer CML (1 year research) | LOW |
| Broadcast technology | Familiar | Vizrt (1 year) | LOW |
| Insurance IT / Business process automation | Familiar | Generali (2 years) | LOW |
| Security / DevSecOps | Proficient | Swisscom Security Champion ×3 | MED |
| Blockchain / Web3 | Familiar | Personal — RPC APIs, basic Solidity, Kraken since 2017 | LOW (bonus only) |
---
## Category 10: Certifications (Skills Signals)
| Certification | Issuer | Year | Active | Resume Weight |
|--------------|--------|------|--------|---------------|
| AWS Certified Solutions Architect Associate | AWS | 2024 | Yes (until Sep 2027) | HIGH |
| Data Engineering with AWS (Nanodegree) | Udacity | 2026 | Yes | HIGH |
| iSAQB Certified Professional for Software Architecture — Foundation Level | iSAQB | 2016 | Yes (no expiry) | MED |
| ITIL® Foundation Certificate in IT Service Management | PEOPLECERT / AXELOS | 2016 | Yes (no expiry) | LOW |
| AI for Trading Nanodegree | Udacity / WorldQuant | 2021 | Yes | LOW (niche) |
| Swisscom Security Champion | Swisscom (internal) | 20232026 | Active | MED (as bullet, not cert line) |
| IBM AI Engineering Specialization | IBM / Coursera | — | Yes | LOW |
---
## Skills Config Guide (for resume generation)
Refers to `config.md` skills layout: **4-3-2-2-2** (resume) or **4-4-3-3-3** (CV).
### Suggested Resume Skills Groups (5 groups)
| Group | Label | Skills to include |
|-------|-------|------------------|
| 1 (4 lines) | Languages & Data | Python, PySpark, SQL (Oracle · Impala · Teradata · Postgres), Java · C# |
| 2 (3 lines) | Cloud & Infra | AWS (S3 · Glue · Athena · Redshift · Airflow), Kubernetes · Docker · Ansible, GitLab CI/CD · Jenkins |
| 3 (2 lines) | Pipelines & Platforms | Kafka · Airflow · SAP BODS · Hadoop, Teradata DWH · ETL/ELT design |
| 4 (2 lines) | ML & Observability | ML inference deployment · MLOps · PyTorch · Scikit-learn, ELK Stack · Grafana · Prometheus |
| 5 (2 lines) | Certifications | AWS Certified Solutions Architect Associate (active), iSAQB CPSA Foundation · ITIL v3 · Data Engineering with AWS (Udacity) |
**Adjust per JD:** For ML/AI roles, swap group 4 to lead with ML; for Platform/Infra roles, expand cloud group. The cert line (group 5) is fixed per `config.md`.
Never add a skill only to mirror a JD. Every listed skill must have a canonical evidence level and an interview-ready example.